Agent skill

Toolkit

by notque in notque/vexjoy-agent

Toolkit management: create and evaluate skills and agents, manage routing tables, generate Claude.md.

MITAuto-check: notesAgent Workflows

Install Toolkit

skills CLI
$ npx skills add notque/vexjoy-agent --skill toolkit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install notque/vexjoy-agent toolkit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/notque/vexjoy-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/meta/toolkit .claude/skills/toolkit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
toolkit
GitHub stars
435
Token cost
~3.1k tokens
SKILL.md length
1,132 words
Files
91 (incl. scripts, references, assets)
Skills in repo
61
Repo updated
First seen
Licence
MIT

At a glance

Toolkit management: create and evaluate skills and agents, manage routing tables, generate Claude.md.

  • Works in 5 steps: Capture intent. What should the skill… → Duplicate check. Run grep -i ""… → Write SKILL.md. Follow… → …
  • Tasks that involve Agent instruction files
  • SKILL.md covers Mode Selection, Create Skill, Create Agent and Uplift for Weaker Models, plus 8 more sections
  • Calls python3

What it does

Toolkit is an agent skill from notque/vexjoy-agent. Toolkit management: create and evaluate skills and agents, manage routing tables, generate Claude.md.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 96 other files, including scripts, reference files and assets (for example `agents/skill-creator/analyzer.md`, `agents/skill-creator/comparator.md` and `agents/skill-creator/grader.md`).

It sits in Agent Workflows, covering Agent instruction files, Building AI agents and Agent evaluation and testing. The repository describes itself as: VexJoy AI Agent with Jev Intelligent Routing - /do routes plain-English requests to the right specialist agent and gates the work with reviews, tests, and a learning loop. The licence is MIT.

When your agent uses it

  • Tasks that involve Agent instruction files
  • Tasks that involve Building AI agents
  • Tasks that involve Agent evaluation and testing

Example prompts

  • “/toolkit”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash, Glob, Grep, Agent, Task, Skill

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Capture intent. What should the skill do? When should it trigger? What output? Are outputs objectively verifiable (code, data) or…
  2. Duplicate check. Run grep -i "" skills/*/SKILL.md to check existing coverage. If an umbrella skill covers the domain, add a reference file…
  3. Write SKILL.md. Follow references/skill-creator/skill-template.md for frontmatter structure. Apply Dense-Complete Writing standard…
  4. Test. Try 3 should-trigger, 2 should-not-trigger, and 2 near-miss prompts with the skill loaded. Revise the SKILL.md until routing and…
  5. Register. Run python3 scripts/generate-skill-index.py to update routing.

What it can do on your machine

Read from SKILL.md and the folder at commit 5218674. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash
    • Glob
    • Grep
    • Agent
    • Task
    • Skill

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Toolkit loads about 3.1k tokens when it runs, and up to ~112k if it reads all its reference files. Until then it costs about 27 tokens; SKILL.md has 1,132 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~112k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash, Glob, Grep, Agent, Task, Skill

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from notque/vexjoy-agent at commit 5218674, republished under its MIT licence (© notque). 1,132 words, ~3,069 tokens.

Download SKILL.mdSave it as .claude/skills/toolkit/SKILL.md (or your agent's skills folder). This skill also uses 90 other files; get the full folder from GitHub.
name
toolkit
description
Toolkit management: create and evaluate skills and agents, manage routing tables, generate Claude.md.
allowed-tools
Read, Write, Edit, Bash, Glob, Grep, Agent, Task, Skill
user-invocable
true
agent
toolkit-governance-engineer
routing.force_route
true
routing.not_for
application code (use workflow), code review (use review)
routing.triggers
create skill, create agent, scaffold skill, scaffold agent, new skill, new agent, skill template, agent template, eval skill, evaluate agent, benchmark skill…
routing.category
meta-tooling
routing.pairs_with
review, workflow

Toolkit

Nine modes covering the full toolkit lifecycle: creating and improving skills and agents; evaluating agents; maintaining routing tables; generating CLAUDE.md; composing multi-skill DAGs; and running the evolution loop. Classify the request and follow the matching section.

Mode Selection

ModeSignalsSection
Skill Creatorcreate skill, scaffold skill, new skill, build a skillCreate Skill
Agent Creatorcreate agent, scaffold agent, new agentCreate Agent
Weak-Model Upliftweaker model, uplift skill, make a skill work for Opus 4.6, improve guidance from generated outputUplift for Weaker Models
Agent Comparisoncompare agents, A/B test agents, benchmark agents, benchmark skill, bake-offCompare Agents
Agent Evaluationevaluate agent quality, audit agent, grade agent, eval skillEvaluate Agent
Skill Composercompose skills, DAG orchestration, skill pipelineCompose Skills
Routing Tablesupdate routing tables, sync routing, routing driftUpdate Routing
Toolkit Evolutionevolve toolkit, self-improve, discover gapsEvolve Toolkit
Generate CLAUDE.mdgenerate claude.md, create claude.md, initGenerate CLAUDE.md

Create Skill

Phases: INTENT -> DRAFT -> TEST -> REGISTER

  1. Capture intent. What should the skill do? When should it trigger? What output? Are outputs objectively verifiable (code, data) or subjective (writing, design)?
  2. Duplicate check. Run grep -i "<domain>" skills/*/SKILL.md to check existing coverage. If an umbrella skill covers the domain, add a reference file instead.
  3. Write SKILL.md. Follow references/skill-creator/skill-template.md for frontmatter structure. Apply Dense-Complete Writing standard. Frontmatter must include: name, description, routing (triggers, not_for, category, pairs_with), allowed-tools.
  4. Test. Try 3 should-trigger, 2 should-not-trigger, and 2 near-miss prompts with the skill loaded. Revise the SKILL.md until routing and output are right.
  5. Register. Run python3 scripts/generate-skill-index.py to update routing.

Load references/skill-creator.md for the full workflow. Deep references in references/skill-creator/ cover progressive disclosure, artifact schemas, complexity tiers, error catalog, enrichment workflow, and more.

Scripts: scripts/skill-creator/


Create Agent

Phases: DISCOVER -> DESIGN -> SCAFFOLD -> REGISTER -> VALIDATE

  1. Discover. Check for domain overlap: grep -i "<domain>" agents/*.md. If an existing agent covers the domain, add a references/ file instead.
  2. Design. Decide role type (reviewer/engineer/orchestrator), allowed tools, complexity, triggers (3-6 specific phrases), pairs_with (verify each exists), reference files, description (intent verb + domain + boundary clause), activation cases.
  3. Scaffold. Write the agent file using references/agent-creator/agent-frontmatter-template.md. Follow docs/PHILOSOPHY.md for operator context structure.
  4. Register. Run python3 scripts/generate-agent-index.py.
  5. Validate. Run python3 scripts/validate-references.py to check reference file integrity. Test activation with the 3+2+2 prompt set.

Load references/agent-creator.md for full phases. Deep references in references/agent-creator/ cover design patterns, frontmatter template, eval design.


Uplift for Weaker Models

Improve a skill, agent, or shared guide until a weaker model produces strong output with it. Load references/weak-model-uplift.md and follow its steps:

  1. Pick the target from data. Query ~/.claude/learning/usage.db and learning.db for heavily used or failing skills.
  2. Build tasks and checks first. 4–8 tasks plus 1–2 held-out tasks; deterministic checks and a yes/no rubric written before any run.
  3. Run the arms. No guidance and current guidance, two samples per task minimum, with python3 scripts/weak_model_run.py.
  4. Score and look. Checks, rubric, your own review of every artifact, optional Jev questions on extracted facts.
  5. Turn failures into rules. Concrete values, before/after examples, runnable checks; delete stale instructions; examples from unrelated products.
  6. Rerun the guided arm and held-out tasks; stop when gains flatten or after three rounds.
  7. Report and ship a per-round table with held-out results, cost, and caveats in the PR body.

Compare Agents

Controlled benchmarks comparing agent variants on identical tasks.

  1. Select variants. Identify the agents to compare (2-4 variants).
  2. Design benchmark. Load references/agent-comparison/benchmark-tasks.md. Select 5-10 representative tasks covering the agent's domain.
  3. Execute. Run each task with each variant. Collect: output quality, token usage, tool calls, time.
  4. Grade. Apply rubric from references/agent-comparison/grading-rubric.md. Score each dimension.
  5. Report. Use references/agent-comparison/report-template.md. Include: methodology, per-task scores, aggregate rankings, cost analysis, recommendation.
  6. Optimize. Load references/agent-comparison/optimize-phase.md to improve the winning variant further.

Load references/agent-comparison.md for the full methodology.


Evaluate Agent

Static structural and standards-compliance grading with a 90-point deterministic scorer.

  1. Read the agent file. Extract frontmatter, body sections, reference files.
  2. Score. Apply rubric from references/agent-evaluation/scoring-rubric.md. Categories: identity (15 pts), expertise (20 pts), routing (15 pts), references (15 pts), workflow (15 pts), standards (10 pts).
  3. Report. Use references/agent-evaluation/report-templates.md. Include: per-category scores, specific findings, improvement recommendations.
  4. Batch mode. For multiple agents: references/agent-evaluation/batch-evaluation.md.

Load references/agent-evaluation.md for the full methodology.


Show full SKILL.md (426 more words)Show less

Compose Skills

DAG-based multi-skill orchestration with dependency resolution.

  1. Define the DAG. List skills in execution order. Identify dependencies (skill B needs output from skill A).
  2. Check compatibility. Load references/skill-composer/compatibility-matrix.md. Verify input/output contracts between skills.
  3. Build the pipeline. Load references/skill-composer/composition-patterns.md for orchestration patterns (serial, parallel, fan-out, conditional).
  4. Execute. Run skills in DAG order. Pass outputs between skills via the defined contracts.
  5. Validate. Check all skills completed. Verify final output meets the composite goal.

Load references/skill-composer.md for the full methodology. See references/skill-composer/examples.md for worked examples.

Scripts: scripts/skill-composer/


Update Routing

5-phase pipeline: SCAN -> EXTRACT -> GENERATE -> UPDATE -> VERIFY.

  1. SCAN. Run python3 scripts/generate-skill-index.py to discover all skills and agents.
  2. EXTRACT. Parse frontmatter from each SKILL.md and agent file. Extract triggers, description, category, complexity.
  3. GENERATE. Build skills/INDEX.json and agents/INDEX.json.
  4. UPDATE. Write index files. PostToolUse hooks auto-regenerate on individual edits; this covers bulk changes and drift.
  5. VERIFY. Compare generated index against discovered files. Report missing entries, conflicts, or stale entries.

Load references/routing-table-updater.md for full phases. Deep references in references/routing-table-updater/ cover routing format, extraction patterns, conflict resolution, batch mode.


Evolve Toolkit

7-phase pipeline: DISCOVER -> DIAGNOSE -> PROPOSE -> CRITIQUE -> BUILD -> VALIDATE -> EVOLVE.

  1. DISCOVER. Audit recent sessions for routing failures, skill gaps, agent weaknesses, user friction.
  2. DIAGNOSE. Load references/toolkit-evolution/diagnose-scripts.md. Run gap analysis scripts. Identify patterns.
  3. PROPOSE. Generate 3-5 improvement proposals with expected impact, effort, risk.
  4. CRITIQUE. Apply multi-perspective review to proposals.
  5. BUILD. Implement the approved proposals using the appropriate mode above (create skill, create agent, etc.).
  6. VALIDATE. Run tests and validators on new/changed components.
  7. EVOLVE. Update evolution history at references/toolkit-evolution/evolution-history.md.

Load references/toolkit-evolution.md for the full pipeline.


Generate CLAUDE.md

4-phase pipeline: SCAN -> DETECT -> GENERATE -> VALIDATE.

  1. SCAN. Check for existing CLAUDE.md. If present, write to CLAUDE.md.generated for comparison. Detect language, framework, build system from repo files.
  2. DETECT. Identify domain enrichment opportunities. Load references/generate-claudemd/examples-and-errors.md for language-specific patterns.
  3. GENERATE. Load template from references/generate-claudemd/CLAUDEMD_TEMPLATE.md. Fill sections: overview, commands, architecture, conventions, testing, deployment.
  4. VALIDATE. Run all documented commands. Verify paths exist. Check for secrets in output.

Optional modes: subdirectory CLAUDE.md for monorepos; minimal mode (overview + commands + architecture only).


Deep References

Load when the task needs detailed schemas, templates, or methodology.

ModeKey References
Skill Creatorreferences/skill-creator.md, references/skill-creator/{skill-template,progressive-disclosure,complexity-tiers,error-catalog,enrichment-workflow}.md
Agent Creatorreferences/agent-creator.md, references/agent-creator/{agent-design-patterns,agent-frontmatter-template,agent-eval-design}.md
Weak-Model Upliftreferences/weak-model-uplift.md
Agent Comparisonreferences/agent-comparison.md, references/agent-comparison/{methodology,grading-rubric,benchmark-tasks,report-template,optimize-phase}.md
Agent Evaluationreferences/agent-evaluation.md, references/agent-evaluation/{scoring-rubric,report-templates,batch-evaluation}.md
Skill Composerreferences/skill-composer.md, references/skill-composer/{compatibility-matrix,composition-patterns,skill-patterns,examples}.md
Routing Tablesreferences/routing-table-updater.md, references/routing-table-updater/{routing-format,extraction-patterns,conflict-resolution,examples}.md
Toolkit Evolutionreferences/toolkit-evolution.md, references/toolkit-evolution/{diagnose-scripts,evolution-history,evolve-preferred-patterns}.md
Generate CLAUDE.mdreferences/generate-claudemd.md, references/generate-claudemd/{CLAUDEMD_TEMPLATE,examples-and-errors}.md

Scripts and Agents

ModeScriptsAgents
Skill Creatorscripts/skill-creator/agents/skill-creator/
Skill Composerscripts/skill-composer/--
Weak-Model Upliftscripts/weak_model_run.py (repo root)--
Routing Tablesscripts/routing-table-updater/--
Agent Comparisonscripts/agent-comparison/--

© notque, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 90 other files (scripts, references, assets) in skills/meta/toolkit of notque/vexjoy-agent.

  • SKILL.md
  • agents/skill-creator/analyzer.md
  • agents/skill-creator/comparator.md
  • agents/skill-creator/grader.md
  • assets/skill-creator/eval_viewer.html
  • references/agent-comparison.md
  • references/agent-comparison/benchmark-tasks.md
  • references/agent-comparison/do-creation-compliance-tasks.json
  • references/agent-comparison/examples-and-errors.md
  • references/agent-comparison/grading-rubric.md
  • references/agent-comparison/methodology.md
  • references/agent-comparison/optimization-guide.md
  • references/agent-comparison/optimization-tasks.example.json
  • references/agent-comparison/optimize-phase.md
  • references/agent-comparison/read-only-ops-short-tasks.json
  • … and 76 more

Open the folder on GitHubat commit 5218674

Compare with similar skills

Toolkit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Toolkit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Toolkit this skillnotque/vexjoy-agent435—~3.1kAutomated safety check: NotesMIT
Microsoft Foundrymicrosoft/GitHub-Copilot-for-Azure2551 repos~6.7kAutomated safety check: PassMIT
Agent Self-Customizationnanocoai/nanoclaw31k1 repos~1.5kAutomated safety check: NotesMIT
Harness Evaltech-leads-club/agent-skills7k—~3.9kAutomated safety check: PassCC-BY-4.0
Agentic Harness Design and ReviewNateBJones-Projects/OB14.7k—~1.8kAutomated safety check: PassCustom licence
Learn From History AuditKonghaYao/peri223—~3.5kAutomated safety check: PassApache-2.0

Similar skills

  • Microsoft Foundry

    microsoft/GitHub-Copilot-for-Azure

    Official

    Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end.

    255 GitHub starsUsed in 1 repo~6.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Self-Customization

    nanocoai/nanoclaw

    A decision tree for an agent changing its own setup: edit memory directly, request approval for packages and MCP servers, and delegate code edits to a builder agent.

    31k GitHub starsUsed in 1 repo~1.5k tokens
    Agent WorkflowsAuto-check: notes
  • Harness Eval

    tech-leads-club/agent-skills

    Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports.

    7k GitHub stars~3.9k tokensUpdated 17 days ago
    Agent WorkflowsAuto-check passed
  • Agentic Harness Design and Review

    NateBJones-Projects/OB1

    Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans.

    4.7k GitHub stars~1.8k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Audits recent agent conversation history and turns repeated failures and successes into testable harness improvement proposals that later audits can check.

    223 GitHub stars~3.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Agents Md Generator

    julianromli/opencode-template

    Generate hierarchical AGENTS.md structures for codebases. An agent skill from julianromli/opencode-template.

    144 GitHub starsUsed in 1 repo~1.4k tokens
    Agent WorkflowsAuto-check: notes

More from notque/vexjoy-agent

All 61 skills in this repo
  • Game Asset Generator

    notque/vexjoy-agent

    Deterministic palette/matrix pixel art (not AI). An agent skill from notque/vexjoy-agent.

    435 GitHub stars~2.3k tokensUpdated 4 days ago
    Auto-check: notes
  • PR Workflow

    notque/vexjoy-agent

    Pull request lifecycle: commit, codex review, sync, review, fix, status, cleanup, and PR mining.

    435 GitHub stars~2.8k tokensUpdated 4 days ago
    Auto-check: notes
  • Architecture Deepening

    notque/vexjoy-agent

    Improve architecture across modules by deepening interfaces.

    435 GitHub stars~3.3k tokensUpdated 4 days ago
    Auto-check: notes
  • Code Quality

    notque/vexjoy-agent

    Code quality: cleanup, linting, formatting, quality gates. An agent skill from notque/vexjoy-agent.

    435 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check: notes
  • Codebase Analyzer

    notque/vexjoy-agent

    Statistical rule discovery from Go codebase patterns. An agent skill from notque/vexjoy-agent.

    435 GitHub stars~2k tokensUpdated 4 days ago
    Auto-check: notes
  • Comment Quality

    notque/vexjoy-agent

    Review and fix temporal references in code comments. An agent skill from notque/vexjoy-agent.

    435 GitHub stars~2k tokensUpdated 4 days ago
    Auto-check: notes

Categories

Questions about Toolkit

What does Toolkit do?

Toolkit management: create and evaluate skills and agents, manage routing tables, generate Claude.md. Toolkit is an agent skill from notque/vexjoy-agent.md.

When should I use Toolkit?

Toolkit fits situations like: tasks that involve Agent instruction files; tasks that involve Building AI agents; tasks that involve Agent evaluation and testing.

How do I install Toolkit in Claude Code?

Run `npx skills add notque/vexjoy-agent --skill toolkit -a claude-code`. Or copy the skill folder (skills/meta/toolkit in notque/vexjoy-agent) into .claude/skills/toolkit in your project. Claude Code loads it when a task matches its description.

How do I install Toolkit in Codex?

Run `npx skills add notque/vexjoy-agent --skill toolkit -a codex`. Or copy the skill folder (skills/meta/toolkit in notque/vexjoy-agent) into .agents/skills/toolkit in your project. Codex loads it when a task matches its description.

Can I use Toolkit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add notque/vexjoy-agent --skill toolkit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/toolkit, .gemini/skills/toolkit, .github/skills/toolkit and .opencode/skills/toolkit in your project.

What does Toolkit need to run?

Going by SKILL.md and its folder, Toolkit needs the command-line tools its instructions call (python3). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, Glob, Grep, Agent, Task, Skill.

Does Toolkit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Toolkit safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Toolkit use?

Toolkit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Toolkit use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 109k tokens, read only when the agent opens those files.

What are the alternatives to Toolkit?

Skills that share tags, products or a category with Toolkit: Microsoft Foundry (microsoft/GitHub-Copilot-for-Azure, 255 stars), Agent Self-Customization (nanocoai/nanoclaw, 31k stars), Harness Eval (tech-leads-club/agent-skills, 7k stars) and Agentic Harness Design and Review (NateBJones-Projects/OB1, 4.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Toolkit?

notque (a GitHub user) maintains it in notque/vexjoy-agent, which has 435 GitHub stars. The repository holds 61 skills in this directory. The repository was last updated on October 3, 2026.

Source: notque/vexjoy-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.