Agent skill

Verbalized Sampling

by gnurio in gnurio/nurijanian-skills

Generate diverse outputs by prompting for a probability distribution instead of a single response.

MITAuto-check: warningsAgent Workflows

Install Verbalized Sampling

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add gnurio/nurijanian-skills --skill verbalized-sampling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gnurio/nurijanian-skills verbalized-sampling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gnurio/nurijanian-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/verbalized-sampling .claude/skills/verbalized-sampling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verbalized-sampling
GitHub stars
124
Token cost
~2k tokens
SKILL.md length
811 words
Files
6 (incl. scripts, references)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Generate diverse outputs by prompting for a probability distribution instead of a single response.

  • Works in 4 steps: Is this naive/obvious? (would it appear… → Is this actionable? (could execution… → Does this require magical thinking?… → …
  • The task needs genuine diversity: creative writing
  • SKILL.md covers Universal VS Template, Variant Selection, Context-First Phase (run… and Critique-and-Improve Loop (run…, plus 6 more sections
  • Runs Python scripts from its folder; calls python

What it does

Verbalized Sampling is an agent skill from gnurio/nurijanian-skills. Generate diverse outputs by prompting for a probability distribution instead of a single response. Implements Verbalized Sampling (VS) from Zhang et al. 2025 — a training-free technique that counteracts LLM mode collapse caused by typicality bias in alignment data. Use when the task needs genuine diversity: creative writing, brainstorming/ideation, synthetic data generation, persona/dialogue simulation, adversarial examples, open-ended QA with multiple valid answers, or any situation where "generate 5 ideas"…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/critique-framework.md`, `references/judges.md` and `references/paper-insights.md`).

It sits in Agent Workflows, covering Brainstorming, Test data and fixtures and Creative writing and fiction. The repository describes itself as: Claude Code and Cursor skills for product managers — PM coaching, verbalized sampling, tech sensemaking, and more. The licence is MIT.

When your agent uses it

  • The task needs genuine diversity: creative writing
  • Brainstorming/ideation
  • Synthetic data generation
  • Persona/dialogue simulation

Example prompts

  • “generate 5 ideas”
  • “/verbalized-sampling”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Is this naive/obvious? (would it appear in a top-10 listicle?)
  2. Is this actionable? (could execution start Monday without further research?)
  3. Does this require magical thinking? (assumes steps will "just work" with no mechanism)
  4. Does this ignore base rates? (approaches with known low success rates presented as good bets)

What it can do on your machine

Read from SKILL.md and the folder at commit 43a0566. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verbalized Sampling loads about 2k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 164 tokens; SKILL.md has 811 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~164
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningTells the agent its actions are pre-authorized / not to stop for confirmationSKILL.md:63
    If you cannot answer Step 2 without asking the user, **ask first** before generating. Generic outputs caused by thin con

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from gnurio/nurijanian-skills at commit 43a0566, republished under its MIT licence (© gnurio). 811 words, ~1,974 tokens.

Download SKILL.mdSave it as .claude/skills/verbalized-sampling/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
verbalized-sampling
description
Generate diverse outputs by prompting for a probability distribution instead of a single response. Implements Verbalized Sampling (VS) from Zhang et al. 2025 — a training-free technique that counteracts LLM mode collapse caused by typicality bias in alignment data. Use when the task needs genuine diversity: creative writing, brainstorming/ideation, synthetic data generation, persona/dialogue simulation, adversarial examples, open-ended QA with multiple valid answers, or any situation where "generate 5 ideas" keeps returning the same cluster. Do NOT use for: single correct answer tasks, factual lookup, strict format compliance.

Verbalized Sampling

Universal VS Template

[Task description with rich context]

Generate {k} responses. Return in JSON format with key "{output_key}" (list of dicts). Each dict:
• text: [output specification]
• probability: estimated probability (0.0–1.0) of this response given the input

{Distribution constraint}

Output ONLY the JSON object.

Distribution constraints — pick one:

  • Sample from the full distribution. — balanced, moderate diversity
  • Sample from the tails of the distribution, with each probability below 0.10. — high diversity
  • Sample from the tails of the distribution, with each probability below 0.01. — maximum diversity

Variant Selection

VariantWhen to useTrade-off
VS-StandardStraightforward tasks, speed priorityBest balance
VS-CoTComplex tasks needing quality + diversitySlight diversity cost, higher quality
VS-MultiMaximum diversity, token cost acceptableBest diversity, 2× token cost

VS-CoT: add "reasoning": "step-by-step thought process" as the first field in each dict.

VS-Multi: Turn 1 generates k/2 responses. Turn 2: "Generate k alternative responses to the original prompt — do not repeat ideas from Turn 1."

Context-First Phase (run before VS)

VS outputs are only as good as the problem framing going in. Before constructing the VS prompt:

Step 1 — Decompose into subproblems: Break the task into 3–5 distinct subproblems or angles. Example: "improve sales for a B2B SaaS" → (1) acquisition channels, (2) conversion from trial, (3) pricing/packaging, (4) referral/word-of-mouth, (5) partnerships.

Step 2 — Load context for each subproblem:

  • What are the real constraints? (time, budget, team size, org politics, market saturation)
  • What do others in this space actually do? (base rates — what approaches are common, what have failed)
  • What has already been tried? (avoid re-suggesting)

Step 3 — Inject context into the VS prompt: Compress answers from Step 2 into the prompt preamble. Name the subproblems as explicit coverage requirements: "Cover at least one idea addressing each of: [subproblem 1], [subproblem 2], ..."

If you cannot answer Step 2 without asking the user, ask first before generating. Generic outputs caused by thin context are the primary failure mode for brainstorming tasks (FM-2).

Critique-and-Improve Loop (run after VS, before presenting)

After generating VS output, run a self-critique pass before presenting results. See references/critique-framework.md for the full 6-dimension framework and prompt templates.

Quick pass: For each output item, check:

  1. Is this naive/obvious? (would it appear in a top-10 listicle?)
  2. Is this actionable? (could execution start Monday without further research?)
  3. Does this require magical thinking? (assumes steps will "just work" with no mechanism)
  4. Does this ignore base rates? (approaches with known low success rates presented as good bets)

If 2+ items fail 2+ checks:

  • Flag the specific failures with reasons
  • Generate 2–3 improved variants that directly address the flagged weaknesses
  • Present: original outputs + critique summary + improved variants

For automated quality scoring of VS outputs, see references/judges.md for LLM-as-Judge prompts.

Output Mode Selection

JSON mode (default for agent pipelines — pipe-able, machine-readable):

  • Use when outputs feed into downstream processing, storage, or evaluation
  • Return raw JSON as-is

Readable mode (in-chat or external sharing):

  • Format using scripts/format_vs_output.py (see below), or render inline as numbered markdown
  • Group by diversity tier: High (p < 0.05), Moderate (p 0.05–0.15), Low/Common (p > 0.15)
  • Show probabilities inline as (p=0.07)

To format manually in-chat:

markdown
## High diversity (p < 0.05)
1. [text] (p=0.03)

## Moderate diversity (p 0.05–0.15)
2. [text] (p=0.08)

CLI formatting: echo '<json>' | python ~/.cursor/skills/verbalized-sampling/scripts/format_vs_output.py

Show full SKILL.md (331 more words)Show less

Probability Threshold Quick Reference

ThresholdUse case
Full distributionGeneral brainstorm, want common + uncommon mix
p < 0.15Moderate novelty — avoids top-5 obvious answers
p < 0.10High diversity — noticeably non-obvious outputs
p < 0.05Aggressive — expect surprising, niche ideas
p < 0.01Maximum — edge cases, stress testing, adversarial

Failure Modes (from empirical evals)

FM-1: Overfit Topic Collapse High-frequency training topics (weight loss, productivity, exercise) resist VS even at p<0.01. The tail of the model's distribution is still inside the well-known solution cluster. The paper's 1.6-2.1× diversity gains apply to creative and niche domains — not saturated self-help topics.

Mitigation: Add explicit exclusion constraints: "Exclude any idea covered in mainstream [domain] journalism. Prioritize ideas from adjacent fields or underrepresented subcultures."

FM-2: Context Starvation → Generic Gravity Thin prompt context ("Xero + retention") produces generic-category outputs even at tail sampling. The more proprietary and specific the context, the better VS performs.

Mitigation: Load rich context before the VS prompt — company stage, current channels, known constraints, target segment, what's already been tried.

FM-3: Semantic Clustering Despite Syntactic Diversity Tail sampling can produce a list that looks different but covers the same solution space. VS does not automatically cross problem-frame boundaries.

Mitigation: Name the problem frames explicitly: "Cover at least one idea from each of: distribution, pricing, community, product, and partnerships."

FM-4: Probability Spread Collapse If the highest and lowest probabilities in your output are within 3× of each other (e.g., all between 0.05–0.09), you're likely in an overfit topic and diversity is illusory.

Diagnosis signal: Good VS output has a spread of at least 5-10× between highest and lowest probability. If spread is tight, switch to FM-1/FM-2 mitigations.

Meta-Prompt: Generate a VS Prompt

I need to generate diverse {output_type} for {use_case}.

Create a Verbalized Sampling prompt that:
1. Clearly describes the task with specific context about {use_case}
2. Requests k={number} outputs in JSON format
3. Requires each output to include "text" and "probability" fields
4. Specifies a distribution constraint appropriate for the diversity level needed:
   - 0.10–0.15 for moderate diversity
   - 0.05–0.10 for high diversity
   - 0.01–0.05 for maximum diversity
5. Ends with "Output ONLY the JSON"
6. Includes explicit problem-frame coverage if topic is likely overfit

Domain Templates

See references/templates.md for ready-to-paste prompts across: creative writing, brainstorming/ideation, dialogue simulation, synthetic data, adversarial examples, open-ended QA.

When NOT to Use VS

  • Single correct answer exists
  • Factual lookup or retrieval
  • Task requires strict format compliance
  • You need one best answer, not a distribution
  • Overfit topic with no rich context available (add context first, then re-evaluate)

© gnurio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/verbalized-sampling of gnurio/nurijanian-skills.

  • SKILL.md
  • references/critique-framework.md
  • references/judges.md
  • references/paper-insights.md
  • references/templates.md
  • scripts/format_vs_output.py

Open the folder on GitHubat commit 43a0566

Compare with similar skills

Verbalized Sampling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verbalized Sampling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verbalized Sampling this skillgnurio/nurijanian-skills124—~2kAutomated safety check: WarnMIT
Creative Thinking For ResearchOrchestra-Research/AI-Research-SKILLs13k3 repos~5.5kAutomated safety check: PassMIT
Light Idea CritiqueLight0305/Light-skills641—~4.7kAutomated safety check: PassMIT
Divergebrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~861Automated safety check: NotesCustom licence
Premise Workshopdanjdewhurst/story-skills2791 repos~2.9kAutomated safety check: NotesMIT
Story IdeatorThomasHoussin/Claude-Book120—~1.6kAutomated safety check: PassMIT

Similar skills

  • Creative Thinking For Research

    Orchestra-Research/AI-Research-SKILLs

    Applies cognitive science frameworks for creative thinking to CS and AI research ideation.

    13k GitHub starsUsed in 3 repos~5.5k tokens
    Agent WorkflowsAuto-check passed
  • Light Idea Critique

    Light0305/Light-skills

    Light 科研主线第 4 步·审 idea:以顶会审稿人标准严审 idea,撞车/无创新 fatal flaw 一票否决(critical 门), 逼出真能发表的 idea。何时用:用户问"这 idea 行不行/够不够新/能不能发""帮我严审/挑刺/找致命问题" / idea 定稿前把关 / 收到 idea-generation 的候选要审 / 怀疑撞车(被人做过)。触发词:审 idea /…

    641 GitHub stars~4.7k tokensUpdated 3 mo ago
    Agent WorkflowsAuto-check passed
  • Diverge

    brycewang-stanford/Auto-Empirical-Research-Skills

    Before implementing, generate 3-5 conceptually distinct approaches labeled by creativity dimension (Novel, Surprising, Diverse, Conventional), then hold for selection.

    4.5k GitHub stars~861 tokensUpdated 2 days ago
    Agent WorkflowsAuto-check: notes
  • Premise Workshop

    danjdewhurst/story-skills

    This skill should be used when the user asks to "brainstorm a story idea", "I have an idea for a story", "what if", "develop a premise", "is this idea strong enough", "workshop my logline"…

    279 GitHub starsUsed in 1 repo~2.9k tokens
    Writing & ContentAuto-check: notes
  • Story Ideator

    ThomasHoussin/Claude-Book

    Generate original storylines from any universe bible without plagiarizing source material.

    120 GitHub stars~1.6k tokensUpdated 8 mo ago
    Agent WorkflowsAuto-check passed
  • Synthesis Build

    acogood/diffmode_free

    Synthesis BUILD stage — the FINAL stage of the Diffmode growth-tactics pipeline (fuses white-space ideation + founder-fit adaptation & merge — formerly two separate synthesis passes — into ONE…

    163 GitHub stars~5.3k tokensUpdated 1 mo ago
    Product & Project ManagementAuto-check passed

More from gnurio/nurijanian-skills

All 15 skills in this repo
  • Vibe Code Leaf Finder

    gnurio/nurijanian-skills

    This skill should be used when a user wants to identify which parts of a codebase are safe to modify with AI ("vibe code") and which require careful human engineering.

    124 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check: notes
  • Corporate Misalignment Finder

    gnurio/nurijanian-skills

    Find, diagnose, and fix misalignment in corporate settings. An agent skill from gnurio/nurijanian-skills.

    124 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Focal Point Finder

    gnurio/nurijanian-skills

    This skill should be used when someone needs to find, propose, or evaluate a focal point in a coordination, negotiation, or alignment problem.

    124 GitHub stars~3.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Orchestrate Pm Alignment

    gnurio/nurijanian-skills

    PM alignment coach and router. An agent skill from gnurio/nurijanian-skills.

    124 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Tech Sensemaking

    gnurio/nurijanian-skills

    Analyze technology announcements to surface non-obvious strategic implications using Verbalized Sampling.

    124 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Adversarial Roleplay

    gnurio/nurijanian-skills

    PM stress-test roleplay. An agent skill from gnurio/nurijanian-skills.

    124 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Verbalized Sampling

What does Verbalized Sampling do?

Generate diverse outputs by prompting for a probability distribution instead of a single response. Verbalized Sampling is an agent skill from gnurio/nurijanian-skills. Generate diverse outputs by prompting for a probability distribution instead of a single response.

When should I use Verbalized Sampling?

Verbalized Sampling fits situations like: the task needs genuine diversity: creative writing; brainstorming/ideation; synthetic data generation; persona/dialogue simulation.

How do I install Verbalized Sampling in Claude Code?

Run `npx skills add gnurio/nurijanian-skills --skill verbalized-sampling -a claude-code`. Or copy the skill folder (skills/verbalized-sampling in gnurio/nurijanian-skills) into .claude/skills/verbalized-sampling in your project. Claude Code loads it when a task matches its description.

How do I install Verbalized Sampling in Codex?

Run `npx skills add gnurio/nurijanian-skills --skill verbalized-sampling -a codex`. Or copy the skill folder (skills/verbalized-sampling in gnurio/nurijanian-skills) into .agents/skills/verbalized-sampling in your project. Codex loads it when a task matches its description.

Can I use Verbalized Sampling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gnurio/nurijanian-skills --skill verbalized-sampling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verbalized-sampling, .gemini/skills/verbalized-sampling, .github/skills/verbalized-sampling and .opencode/skills/verbalized-sampling in your project.

What does Verbalized Sampling need to run?

Going by SKILL.md and its folder, Verbalized Sampling needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Verbalized Sampling access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Verbalized Sampling safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Verbalized Sampling use?

Verbalized Sampling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verbalized Sampling use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.7k tokens, read only when the agent opens those files.

What are the alternatives to Verbalized Sampling?

Skills that share tags, products or a category with Verbalized Sampling: Creative Thinking For Research (Orchestra-Research/AI-Research-SKILLs, 13k stars), Light Idea Critique (Light0305/Light-skills, 641 stars), Diverge (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars) and Premise Workshop (danjdewhurst/story-skills, 279 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verbalized Sampling?

gnurio (a GitHub user) maintains it in gnurio/nurijanian-skills, which has 124 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on August 13, 2026.

Source: gnurio/nurijanian-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.