Agent skill

Grill Jev

by notque in notque/vexjoy-agent

Broad Jev interrogation: generate up to 50 context-specific questions about a plan, spec, design, code artifact, or any topic — using all primitive shapes and structured forms to surface hidden…

MITAuto-check: notes

Install Grill Jev

skills CLI
$ npx skills add notque/vexjoy-agent --skill grill-jev -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install notque/vexjoy-agent grill-jev --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/notque/vexjoy-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/meta/grill-jev .claude/skills/grill-jev && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
grill-jev
GitHub stars
438
Token cost
~3.6k tokens
SKILL.md length
1,206 words
Files
2 (incl. references)
Skills in repo
61
Repo updated
First seen
Licence
MIT

At a glance

Broad Jev interrogation: generate up to 50 context-specific questions about a plan, spec, design, code artifact, or any topic — using all primitive shapes and structured forms to surface hidden…

  • Works in 4 steps: Read the artifact — the plan, spec,… → Generate questions — the LLM executing… → Send to Jev — evaluate the questions… → …
  • SKILL.md covers When to invoke, How it works, Asking good questions and Question categories, plus 6 more sections
  • Calls python3

What it does

Grill Jev is an agent skill from notque/vexjoy-agent. Broad Jev interrogation: generate up to 50 context-specific questions about a plan, spec, design, code artifact, or any topic — using all primitive shapes and structured forms to surface hidden problems before they ship.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/question-battery.md`).

The repository describes itself as: VexJoy AI Agent with Jev Intelligent Routing - /do routes plain-English requests to the right specialist agent and gates the work with reviews, tests, and a learning loop. The licence is MIT.

Example prompts

  • “/grill-jev”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash, Glob, Grep

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Read the artifact — the plan, spec, design, or code being interrogated
  2. Generate questions — the LLM executing this skill produces up to 50 context-specific Jev questions tailored to the artifact. It writes the…
  3. Send to Jev — evaluate the questions against the artifact as state. Use Vercel AI Gateway. Too much context is the most common failure…
  4. Report findings — high-signal answers with suggested actions

What it can do on your machine

Read from SKILL.md and the folder at commit 5218674. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash
    • Glob
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Grill Jev loads about 3.6k tokens when it runs, and up to ~6.2k if it reads all its reference files. Until then it costs about 58 tokens; SKILL.md has 1,206 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash, Glob, Grep

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from notque/vexjoy-agent at commit 5218674, republished under its MIT licence (© notque). 1,206 words, ~3,604 tokens.

Download SKILL.mdSave it as .claude/skills/grill-jev/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
grill-jev
description
Broad Jev interrogation: generate up to 50 context-specific questions about a plan, spec, design, code artifact, or any topic — using all primitive shapes and structured forms to surface hidden problems before they ship.
allowed-tools
Read, Write, Edit, Bash, Glob, Grep
user-invocable
true
routing.force_route
true
routing.triggers
grill, validate plan, validate spec, validate design, interrogate, stress test plan, stress test spec, 50 questions, deep validation, grill this, grill the…
routing.not_for
Writing new Jev programs (use building-with-jev). Routing requests (use do). Code review for quality/style (use review). Security scanning (use security)…
routing.pairs_with
building-with-jev, review
routing.complexity
Complex
routing.category
meta

Grill-Jev

Generate a context-specific Jev interrogation battery against a plan, spec, design, code artifact, or any topic. Questions are generated from the artifact itself — not from a static list — so they target the actual risks and gaps in what you provide.

When to invoke

  • User says "grill this", "validate my plan", "find holes in my spec", "sanity check", "stress test this", "50 questions", "interrogate", "grill-jev"
  • Automatically after any planning output. When a plan, spec, or design is produced — before execution begins — pass it through grill-jev. Any structured output with phases, steps, or checklist items qualifies.
  • When the user says "does this look right", "approve this", "is this ready", "review my plan"

How it works

  1. Read the artifact — the plan, spec, design, or code being interrogated
  2. Generate questions — the LLM executing this skill produces up to 50 context-specific Jev questions tailored to the artifact. It writes the JSON battery to a temporary file; no secondary model call or API key is used. Questions use all structured shapes:
    • Noul with {true: {what, examples}, false: {what, examples}} criteria
    • Choice with {what, not_for, examples} per option
    • Score with {summary, signals} per level
    • Noul with array compare instructions when two state paths need comparison
  3. Send to Jev — evaluate the questions against the artifact as state. Use Vercel AI Gateway. Too much context is the most common failure: the artifact is resent with every batch, so split the battery by estimated tokens (jev_limits.request_tokens), not by question count, keeping each request under the size target in skills/shared-patterns/jev-production-lessons.md. When the artifact alone passes the target, send the sections each question needs instead of the whole artifact. scripts/grill-jev.py does the splitting: it packs batches by tokens, and when a batch fails it halves and retries down to single questions, so one oversized or malformed question costs only its own answer. A question that still fails is listed as UNANSWERED (unknown, never a pass) and the run exits 2. If a tiny probe also fails, Jev is down and the run stops. The script rejects a malformed battery before sending (for example, Score levels must be a list).
  4. Report findings — high-signal answers with suggested actions

Asking good questions

Question quality depends on context and framing. The rules below were validated through 3 iterative Jev loops.

Context inputs that improve questions:

InputWhy it mattersExample
AudienceWho executes or approves this? SRE, junior dev, product owner, external team?SREs need ops-specific questions; product owners need outcome and risk questions
System contextWhat does the system actually do? What are its constraints?Stateful vs. stateless changes need different risk questions
PurposeWhat decision does this artifact support?Approval gate → binary questions; exploration → broader coverage
Depth wanted15 sharp questions or 50 comprehensive ones?Match count to stakes and complexity

Pass context via --context "audience: SRE, system: stateful payment service, known constraint: cannot have >5min downtime".

Generation rules (Jev-validated):

  1. Specific over generic — questions must name specific steps, systems, or claims in the artifact. "Does step 4's migration define a rollback safe to run under live traffic?" beats "does this have rollback?".

  2. Mentions trigger deeper scrutiny, not shallower — when the artifact mentions a risk, gap, or uncertainty, generate MORE targeted questions about it. Acknowledgment is not mitigation. "This is risky" without a defined mitigation is itself a finding.

  3. Audience weight — use audience context to focus questions. An SRE needs ops questions. A junior dev needs step-clarity questions. A product owner needs outcome questions.

  4. Coverage balance — aim for breadth across relevant categories (completeness, feasibility, risk, scope, verification, consistency, reversibility, security, cost and throughput). Do not cluster all questions on one category. Skip a category only when the artifact has nothing that triggers it.

  5. Scale by complexity — simple artifact (15-20 questions), medium (25-35), complex (40-50). Hard cap: 50.

Self-calibration loop — if findings feel generic or off-target, use Jev to improve:

  1. Run grill-jev — observe which findings feel shallow
  2. Ask Jev: "Which questions were not specific to this artifact? What context would have produced better questions?"
  3. Feed that context back via --context and re-run
Show full SKILL.md (527 more words)Show less

Question categories

Generate questions covering these nine areas, weighted by what the artifact contains:

CategoryWhat it finds
Completenessmissing phases, undefined terms, unstated assumptions
Feasibilityresource constraints, timeline, dependencies
Risk & failure modeswhat happens when each step fails
Scope & boundarieswhat's in/out, integration surfaces
Verificationhow do we know it worked, success criteria
Consistencyinternal contradictions, duplicate effort
Reversibilitycan we undo this, migration risk
Security & safetyauth, data exposure, destructive operations
Cost & throughputcalls, tokens, and requests per run and per second against the provider's documented rate limits; fan-out size; concurrency; retry policy; eval cost

Plans that call Jev or another metered API

A plan can be complete, feasible, and safe and still fail in production because one run spends the provider's per-second limit. Grill it on arithmetic, not just on prose:

  1. Check it against the rules. Load skills/shared-patterns/jev-production-lessons.md and turn every unticked pre-ship checklist item into a cost_ Noul with report_when: "false". Name the exact number in the question: "Does the plan keep every request at or under 4k tokens or a measured reliable size?", "Does it send each stage's requests at once, with an instance cap near floor(0.25 × 250,000 / tokens_per_request)?", "Does it set attempts, per-attempt timeout, run deadline, and a retry budget near 4 × requests × failure rate?"
  2. Price it first. When the artifact calls Jev, build (or ask for) a JSON of every request one run sends and run python3 scripts/jev-budget-check.py --payload run.json --concurrency C --concurrent-runs N --attempts A, plus --eval-cases N for any eval the plan runs. A fail is a high-signal finding on its own; a warn goes in the report. Put the check's summary in state under budget so battery questions can inspect it.
  3. Ask about throughput explicitly. Include Cost & throughput questions (see references/question-battery.md): does the plan state tokens per run and per second against the documented limits (250,000 input tokens per second and 1,200 requests per minute for Jev on 2026-09-22)? Does it fan full detail out over every unit, or cascade? Is in-flight concurrency capped? Do retries use jittered exponential backoff with a per-run budget? Is the eval priced and paced? Which errors mean "back off" on the production transport (Vercel AI Gateway reports upstream overload as 503)?
  4. Treat "each request fits" as unproven. A plan that shows every request under the per-request limit has not shown the run fits. Look for the per-second number.
  5. Diagnoses need measurements. When the artifact explains a failure (an outage, a size cap, a bad payload), ask whether it measured the run's own rate and retry count and tested a small known-good request on the same transport before concluding.

Question generation prompt

Use this system prompt to generate the question battery:

You are generating a Jev question battery to interrogate an artifact.
The artifact is provided as state at key "artifact".

Generate between 20 and 50 questions. Scale the count to the artifact's
complexity — a 3-step bug fix needs fewer questions than a multi-service
migration plan.

For each question, choose the most appropriate Jev primitive:
- Noul: yes/no probability. Use when you want to know whether something
  is true or absent. Always add structured criteria (true/false with
  what + examples) when the boundary is non-obvious.
- Choice: pick one from a known set. Use for risk levels, categories,
  reversibility classifications.
- Score: position on a spectrum. Use for completeness, timeline
  realism, detection speed.

Use array compare instructions when two or more fields in the state
need to be compared side by side.

Focus questions on the actual content of the artifact. A plan with no
database steps needs no database migration questions. A plan with a
single deploy step needs no multi-service blast radius question.

Output a JSON dict of question_id -> question definition.

Running the battery

bash
# File artifact
python3 scripts/grill-jev.py --file task_plan.md --mode plan

# Inline text
python3 scripts/grill-jev.py --text "$(cat task_plan.md)" --mode plan

# The executing LLM writes /tmp/grill-questions.json, then Jev evaluates it.
python3 scripts/grill-jev.py --file design.md --questions-file /tmp/grill-questions.json

The executing LLM owns question generation; the script only evaluates a supplied battery through Jev. Without --questions-file, it uses the static fallback battery. The compact guide in references/question-battery.md provides:

  • Example shapes for the LLM question generator
  • Fallback when generation is unavailable
  • Test fixture for unit tests

Output

GRILL-JEV FINDINGS — mode: plan — 31 questions (generated)
============================================================
HIGH SIGNAL (requires attention):
  [completeness/has_success_criteria] noul=0.89 TRUE — no measurable success criteria
    → add explicit success criteria: observable outcomes, not "it works"
  [risk/migration_live_safety] noul=0.84 TRUE — migration runs against live traffic
    → add maintenance window or use online migration tool

CATEGORY SUMMARY:
  completeness  2 findings   risk  1 finding

OVERALL READINESS: 1.4/3 — partially complete; address findings before executing

Primitive shapes — quick reference

python
# Noul with structured criteria
"has_success_criteria": {
    "type": "noul",
    "instructions": {"question": "Does the plan define measurable success criteria?",
                     "inspect": "artifact"},
    "criteria": {
        "true":  {"what": "Specific, observable outcomes are named",
                  "examples": ["all tests pass", "p95 latency < 200ms"]},
        "false": {"what": "Success is vague or absent",
                  "examples": ["it works", "done"]}
    }
}

# Choice with what/not_for/examples
"overall_risk": {
    "type": "choice",
    "instructions": {"question": "What is the overall risk level?",
                     "focus": "Weigh irreversibility, dependency count, blast radius."},
    "criteria": {
        "low":  {"what": "Reversible, few dependencies, limited blast radius",
                 "not_for": "Any step that cannot be undone",
                 "examples": ["adding an optional config flag"]},
        "medium": {"what": "Some irreversibility or cross-system dependencies",
                   "examples": ["schema migration with rollback plan"]},
        "high": {"what": "Irreversible steps, wide blast radius",
                 "examples": ["deleting a table", "replacing auth system"]}
    }
}

# Score with summary/signals levels
"completeness": {
    "type": "score",
    "instructions": {"question": "How complete is this plan?",
                     "note": "Judge whether a competent engineer could execute it without guessing."},
    "criteria": [
        {"summary": "Critically incomplete",
         "signals": ["missing phases", "undefined terms", "no rollback"]},
        {"summary": "Partially complete",
         "signals": ["main path clear", "some steps vague"]},
        {"summary": "Mostly complete",
         "signals": ["all phases named", "minor details missing"]},
        {"summary": "Complete",
         "signals": ["all steps actionable", "success criteria defined"]}
    ]
}

# Noul with array compare instructions
"assumption_vs_reality": {
    "type": "noul",
    "instructions": {
        "question": "Do the plan's assumptions conflict with the known system context?",
        "compare": ["artifact", "context"],
        "focus": "Look for things the plan takes for granted that could be false."
    }
}

Question-authoring guide: references/question-battery.md.

Integration into planning workflows

Any skill or agent that produces a plan should call grill-jev before declaring it complete:

python
import subprocess

result = subprocess.run(
    ["python3", "scripts/grill-jev.py", "--file", plan_path, "--mode", "plan", "--questions-file", questions_path],
    capture_output=True, text=True
)
print(result.stdout)
# Non-zero exit when HIGH SIGNAL findings exceed threshold

© notque, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/meta/grill-jev of notque/vexjoy-agent.

  • SKILL.md
  • references/question-battery.md

Open the folder on GitHubat commit 5218674

Compare with similar skills

Grill Jev next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Grill Jev compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Grill Jev this skillnotque/vexjoy-agent438—~3.6kAutomated safety check: NotesMIT
Grillasgeirtj/system_prompts_leaks69k—~1.7kAutomated safety check: PassMIT
Grill Menrwl/nx29k—~1.2kAutomated safety check: PassMIT
Grillingvinvcn/mattpocock-skills-zh-CN4.7k—~228Automated safety check: PassMIT
Jev Socialsickn33/agentic-awesome-skills47k1 repos~3.4kAutomated safety check: PassMIT
Grill Mefeiskyer/claude-code-settings1.7k—~913Automated safety check: PassMIT

Similar skills

  • Grill

    asgeirtj/system_prompts_leaks

    Run an explicitly requested decision interview and record each settled decision in durable project documentation.

    69k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Grill Me

    nrwl/nx

    Grill the user relentlessly about a plan, design, decision, or set of review findings — working the decision tree in rounds until nothing is left silently assumed.

    29k GitHub stars~1.2k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Grilling

    vinvcn/mattpocock-skills-zh-CN

    围绕计划、decision 或 idea 持续追问用户。适用于用户想对自己的思路做压力测试,或使用任何 “grill” 触发措辞时。

    4.7k GitHub stars~228 tokensUpdated 10 days ago
    Agent WorkflowsAuto-check passed
  • Jev Social

    sickn33/agentic-awesome-skills

    Run read-only, browser-grounded Instagram, TikTok, or LinkedIn research through Jev routing and socai CLI, returning source-linked evidence and reports.

    47k GitHub starsUsed in 1 repo~3.4k tokens
    Auto-check passed
  • Grill Me

    feiskyer/claude-code-settings

    针对方案或设计的高强度追问式面试(adversarial design review / grill session),暴露假设漏洞与缺失约束,过程中同步维护领域模型(术语表和 ADR)。手动调用 /grill-me。

    1.7k GitHub stars~913 tokensUpdated 10 days ago
    Agent WorkflowsAuto-check passed
  • Grill With Docs

    alirezarezvani/claude-skills

    Docs-anchored grilling session — challenges a plan against the project's existing language (CONTEXT.md) and recorded decisions (docs/adr/), and updates those files inline as terminology and…

    28k GitHub stars~1.8k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed

More from notque/vexjoy-agent

All 61 skills in this repo
  • Game Asset Generator

    notque/vexjoy-agent

    Deterministic palette/matrix pixel art (not AI). An agent skill from notque/vexjoy-agent.

    438 GitHub stars~2.3k tokensUpdated 5 days ago
    Auto-check: notes
  • PR Workflow

    notque/vexjoy-agent

    Pull request lifecycle: commit, codex review, sync, review, fix, status, cleanup, and PR mining.

    438 GitHub stars~2.8k tokensUpdated 5 days ago
    Auto-check: notes
  • Architecture Deepening

    notque/vexjoy-agent

    Improve architecture across modules by deepening interfaces.

    438 GitHub stars~3.3k tokensUpdated 5 days ago
    Auto-check: notes
  • Code Quality

    notque/vexjoy-agent

    Code quality: cleanup, linting, formatting, quality gates. An agent skill from notque/vexjoy-agent.

    438 GitHub stars~1.5k tokensUpdated 5 days ago
    Auto-check: notes
  • Codebase Analyzer

    notque/vexjoy-agent

    Statistical rule discovery from Go codebase patterns. An agent skill from notque/vexjoy-agent.

    438 GitHub stars~2k tokensUpdated 5 days ago
    Auto-check: notes
  • Comment Quality

    notque/vexjoy-agent

    Review and fix temporal references in code comments. An agent skill from notque/vexjoy-agent.

    438 GitHub stars~2k tokensUpdated 5 days ago
    Auto-check: notes

Questions about Grill Jev

What does Grill Jev do?

Broad Jev interrogation: generate up to 50 context-specific questions about a plan, spec, design, code artifact, or any topic — using all primitive shapes and structured forms to surface hidden…. Grill Jev is an agent skill from notque/vexjoy-agent. Broad Jev interrogation: generate up to 50 context-specific questions about a plan, spec, design, code artifact, or any topic — using all primitive shapes and structured forms to surface hidden problems before they ship.

How do I install Grill Jev in Claude Code?

Run `npx skills add notque/vexjoy-agent --skill grill-jev -a claude-code`. Or copy the skill folder (skills/meta/grill-jev in notque/vexjoy-agent) into .claude/skills/grill-jev in your project. Claude Code loads it when a task matches its description.

How do I install Grill Jev in Codex?

Run `npx skills add notque/vexjoy-agent --skill grill-jev -a codex`. Or copy the skill folder (skills/meta/grill-jev in notque/vexjoy-agent) into .agents/skills/grill-jev in your project. Codex loads it when a task matches its description.

Can I use Grill Jev in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add notque/vexjoy-agent --skill grill-jev -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/grill-jev, .gemini/skills/grill-jev, .github/skills/grill-jev and .opencode/skills/grill-jev in your project.

What does Grill Jev need to run?

Going by SKILL.md and its folder, Grill Jev needs the command-line tools its instructions call (python3). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, Glob, Grep.

Does Grill Jev access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Grill Jev safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Grill Jev use?

Grill Jev is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Grill Jev use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.

What are the alternatives to Grill Jev?

Skills that share tags, products or a category with Grill Jev: Grill (asgeirtj/system_prompts_leaks, 69k stars), Grill Me (nrwl/nx, 29k stars), Grilling (vinvcn/mattpocock-skills-zh-CN, 4.7k stars) and Jev Social (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Grill Jev?

notque (a GitHub user) maintains it in notque/vexjoy-agent, which has 438 GitHub stars. The repository holds 61 skills in this directory. The repository was last updated on October 3, 2026.

Source: notque/vexjoy-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.