Caveman Optimization Evaluator
JuliusBrussee/caveman
Turns a Caveman report-only optimization observation into one minimal code change and a paired baseline evaluation, after the operator picks which to pursue.
This skill helps an LLM generate correct AxAgent tuning and evaluation code using @ax-llm/ax.
$ npx skills add dosco/aithy --skill ax-agent-optimize -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install dosco/aithy ax-agent-optimize --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ax-agent-optimize .claude/skills/ax-agent-optimize && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ax-agent-optimize" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-agent-optimize into .claude/skills/ax-agent-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-agent-optimize", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/dosco/aithy/tree/main/.claude/skills/ax-agent-optimizeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add dosco/aithy --skill ax-agent-optimize -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install dosco/aithy ax-agent-optimize --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/ax-agent-optimize .agents/skills/ax-agent-optimize && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ax-agent-optimize" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-agent-optimize into .agents/skills/ax-agent-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-agent-optimize", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dosco/aithy --skill ax-agent-optimize -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install dosco/aithy ax-agent-optimize --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/ax-agent-optimize .cursor/skills/ax-agent-optimize && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ax-agent-optimize" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-agent-optimize into .cursor/skills/ax-agent-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-agent-optimize", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/dosco/aithy.git --path .claude/skills/ax-agent-optimize--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add dosco/aithy --skill ax-agent-optimize -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install dosco/aithy ax-agent-optimize --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/ax-agent-optimize .gemini/skills/ax-agent-optimize && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ax-agent-optimize" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-agent-optimize into .gemini/skills/ax-agent-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-agent-optimize", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install dosco/aithy ax-agent-optimizeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add dosco/aithy --skill ax-agent-optimize -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/ax-agent-optimize .github/skills/ax-agent-optimize && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ax-agent-optimize" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-agent-optimize into .github/skills/ax-agent-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-agent-optimize", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dosco/aithy --skill ax-agent-optimize -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install dosco/aithy ax-agent-optimize --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/ax-agent-optimize .opencode/skills/ax-agent-optimize && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ax-agent-optimize" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-agent-optimize into .opencode/skills/ax-agent-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-agent-optimize", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ax-agent-optimizeThis skill helps an LLM generate correct AxAgent tuning and evaluation code using @ax-llm/ax.
Ax Agent Optimize is an agent skill from dosco/aithy. This skill helps an LLM generate correct AxAgent tuning and evaluation code using @ax-llm/ax. Use when the user asks about agent.optimize(...), judgeOptions, eval datasets, optimization targets, saved optimizedProgram artifacts, or agent optimization guidance.
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: A personal AI agent that can work safely on your machine, remember useful context, and keep its data under your control. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 0c9855f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
raw.githubusercontent.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ax Agent Optimize loads about 4.8k tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 2,110 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from dosco/aithy at commit 0c9855f, republished under its Apache-2.0 licence (© dosco). 2,110 words, ~4,849 tokens.
.claude/skills/ax-agent-optimize/SKILL.md (or your agent's skills folder).Use this skill for agent.optimize(...) workflows. Prefer short, modern, copyable patterns. Do not repeat general agent-authoring guidance unless the user needs it. For generic ax(...) or flow(...) tuning with top-level optimize(...), use the ax-gepa skill instead.
Your job is to help the model choose a good optimization setup for the user's actual goal:
agent.optimize(...) only after the agent is already configured and runnable.input and criteria, then let agent.optimize(...) use its default actor target and judge-based metric.optimize(program, train, metric, options) for non-agent generators and flows; do not rewrite normal agent task-record examples to the generic helper.metric only when success is easy to score from the prediction and task record.judgeAI plus judgeOptions when the judge should run on a stronger or separate model than the agent runtime model.AxGen evaluator when the user needs LLM-as-judge behavior outside the built-in agent.optimize(...) flow.target unless the user clearly wants responder-only tuning or explicit program IDs.f.object(...) over vague f.json(...) whenever the agent must reason about returned fields.axSerializeOptimizedProgram(result.optimizedProgram!), then restore with axDeserializeOptimizedProgram(saved) and agent.applyOptimization(...).bootstrap is enabled, bootstrapped demos are persisted inside result.optimizedProgram.demos; raw failed traces are not saved in v1.autoUpgrade) appear in captured traces/demos as their truncated preview string, not the full value — same as declared truncate-style contextFields. This is expected; do not treat the shortened value as a bug in the saved demos.train and validation unless the user already has a holdout set.agent.optimize(...) now optimizes generic components exposed by the selected target programs; target: 'actor' only tunes actor components, target: 'responder' only tunes responder components, and target: 'all' broadens the component set.result.optimizedProgram.componentMap is the canonical saved artifact for agent GEPA runs. It may include actor instructions, descriptions, tool descriptions/names, templates, or runtime primitives depending on what the selected target exposes.Pick the optimization shape from the user's need:
expectedActions and forbiddenActions.target: 'responder', but only if the task is not mostly tool-selection or clarification behavior.agent.playbook().evolve(dataset)) to mine failures into verified playbook bullets under a held-out gate. Python, Java, C++, Go, and Rust expose the same loop with native method casing; see ax-playbook. optimize(...) maximizes a metric by tuning instructions and demos; playbook evolution grows durable rules.Choose task design carefully:
Optimization works much better when the agent and dataset remove avoidable ambiguity:
maxSubAgentCalls small in examples unless the user is explicitly testing broad fan-out behavior.javascript: prefixes, mixed prose/code, and multi-snippet turns.Good pattern:
Bad pattern:
json with an underspecified shapeAtlas without clarifying whether that is a project, team, or accountChoose the scoring path based on how objectively the task can be measured:
metric when you can score success directly from prediction and example.judgeOptions.description to tell the built-in judge what to value most.agent.optimize(...) and still wants LLM judging.Quick rules:
AxGen evaluator.Important:
metric overrides the built-in judge path entirely.AxGen.metric and judge guidance unless the user explicitly wants two separate scoring systems and understands only the custom metric drives optimization.AxGen judge metric, prefer a numeric score:number output over a string tier when possible. It is simpler and less fragile in practice.import {
AxAIGoogleGeminiModel,
AxJSRuntime,
axDefaultOptimizerLogger,
agent,
ai,
f,
fn,
axDeserializeOptimizedProgram,
axSerializeOptimizedProgram,
} from '@ax-llm/ax';
const tools = [
fn('sendEmail')
.namespace('email')
.description('Send an email message')
.arg('to', f.string('Recipient email address'))
.arg('body', f.string('Email body text'))
.returns(
f.object({
sent: f.boolean('Whether the email was sent'),
to: f.string('Recipient email address'),
})
)
.handler(async ({ to }) => ({ sent: true, to }))
.build(),
];
const studentAI = ai({
name: 'google-gemini',
apiKey: process.env.GOOGLE_APIKEY!,
config: { model: AxAIGoogleGeminiModel.Gemini31FlashLite, temperature: 0.2 },
});
const judgeAI = ai({
name: 'google-gemini',
apiKey: process.env.GOOGLE_APIKEY!,
config: { model: AxAIGoogleGeminiModel.Gemini35Flash, temperature: 1.0 },
});
const assistant = agent('query:string -> answer:string', {
ai: studentAI,
judgeAI,
contextFields: [],
runtime: new AxJSRuntime(),
functions: tools,
contextPolicy: { preset: 'checkpointed', budget: 'balanced' },
judgeOptions: {
description: 'Prefer correct tool use over polished wording.',
model: 'judge-model',
},
});
const tasks = [
{
input: { query: 'Send an email to Jim saying good morning.' },
criteria: 'Use the email tool and send the message to Jim.',
expectedActions: ['email.sendEmail'],
},
];
const result = await assistant.optimize(tasks, {
maxMetricCalls: 12,
verbose: true,
});
const saved = axSerializeOptimizedProgram(result.optimizedProgram!);
const restored = axDeserializeOptimizedProgram(saved);
assistant.applyOptimization(restored);Start here unless the user clearly needs a hand-built scorer:
const tasks = [
{
input: { query: 'Send an email to Jim saying good morning.' },
criteria: 'Use the email tool and send the message to Jim.',
expectedActions: ['email.sendEmail'],
},
];
const result = await assistant.optimize(tasks);
assistant.applyOptimization(result.optimizedProgram!);target defaults to actor optimization.metric defaults to the built-in LLM judge.judgeAI is optional; if omitted, the agent falls back to its configured judge model or runtime model.bootstrap: true is a good next step for tool-heavy agents when you want GEPA to start from successful traces from the provided tasks.criteria.Use this when the task has crisp correctness and cost/behavior tradeoffs:
const result = await assistant.optimize(tasks, {
target: 'actor',
metric: ({ prediction, example }) => {
if (prediction.completionType !== 'final' || !prediction.output) {
return 0;
}
let score = 0;
if (prediction.output.answer.includes('Jim')) score += 0.4;
if (
prediction.functionCalls.some(
(call) => call.qualifiedName === 'email.sendEmail'
)
) {
score += 0.4;
}
if (prediction.turnCount <= 3) {
score += 0.2;
}
return score;
},
});Use this pattern when:
Use this when the agent behavior needs holistic review:
const result = await assistant.optimize(tasks, {
judgeAI,
judgeOptions: {
model: AxAIGoogleGeminiModel.Gemini35Flash,
description:
'Be strict about unnecessary child-agent calls, weak clarifications, and incorrect tool choices.',
},
maxMetricCalls: 12,
});Use this pattern when:
AxGen Judge PatternUse this only when the user needs LLM judging outside the built-in agent.optimize(...) path:
import { AxGen, s } from '@ax-llm/ax';
const judgeGen = new AxGen(
s(`
taskInput:json "Task input",
candidateOutput:json "Candidate output",
expectedOutput?:json "Optional reference output"
->
score:number "Normalized score from 0 to 1"
`)
);
judgeGen.setInstruction(
'Score the candidate output from 0 to 1. Reward correctness and task completion. Return only the score field.'
);
const metric = async ({ prediction, example }) => {
const result = await judgeGen.forward(judgeAI, {
taskInput: example,
candidateOutput: prediction,
expectedOutput: example.expectedOutput,
});
return Math.max(0, Math.min(1, result.score));
};
const result = await optimizer.compile(program, train, metric, {
validationExamples: validation,
});Use this pattern when:
AxGen, flow, or another program directlyagent.optimize(...) wrapperexpectedActions and forbiddenActions when tool correctness matters.judgeOptions mirrors normal forward options and supports extra judge guidance through description.metric, that overrides the built-in judge path.Decision rules:
AxGen evaluator when the user is not calling agent.optimize(...) but still wants LLM judging.judgeOptions.description to steer the judge toward the user's real priority, such as tool correctness, brevity, groundedness, or policy compliance.mcpEvaluation: 'live' is explicit.ax-mcp for recording/replay transport setup and MCP side-effect policy.AxMCPRecordingTransport to capture a real session once and AxMCPReplayTransport for deterministic optimization/evaluation.AxEventRuntime; do not leave a live subscription active in a default
optimization run.agent.optimize(...) runs each evaluation rollout from a clean continuation state.getState() and setState(...) is not used during eval rollouts.askClarification(...) is treated as a scored evaluation outcome instead of going through the responder.prediction.completionType === 'askClarification', populated prediction.clarification, and absent prediction.output.prediction.completionType === 'final' and populated prediction.output.target: 'responder' still works, but clarification-heavy tasks are usually low-signal for responder optimization.functions: [...] for specialist delegation. Their calls appear as normal function-call records.team.writer(...) only after narrowing tool output in JS."result.optimizedProgram if the user wants portable artifacts.new AxOptimizedProgramImpl(...), then call agent.applyOptimization(...).componentMap reapplies the learned strings.AxGen.json tool returns when the agent must reason about specific fields across tool or child-agent calls.© dosco, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/ax-agent-optimize of dosco/aithy.
Open the folder on GitHubat commit 0c9855f
Ax Agent Optimize next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ax Agent Optimize this skilldosco/aithy | 107 | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | |
| Caveman Optimization EvaluatorJuliusBrussee/caveman | 110k | 1 repos | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| SQL Optimizationgithub/awesome-copilot | 40k | 2 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Database Optimizerdavila7/claude-code-templates | 32k | 8 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Arize Evaluatorgithub/awesome-copilot | 40k | 2 repos | ~8.1k | Automated safety check: Notes | MIT | |
| Postgresql Optimizationdavila7/claude-code-templates | 32k | 4 repos | ~951 | Automated safety check: Pass | MIT |
JuliusBrussee/caveman
Turns a Caveman report-only optimization observation into one minimal code change and a paired baseline evaluation, after the operator picks which to pursue.
github/awesome-copilot
Universal SQL performance optimization assistant for comprehensive query tuning, indexing strategies, and database performance analysis across all SQL databases (MySQL, PostgreSQL, SQL Server…
davila7/claude-code-templates
Expert database optimizer specializing in modern performance tuning, query optimization, and scalable architectures.
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
davila7/claude-code-templates
PostgreSQL database optimization workflow for query tuning, indexing strategies, performance analysis, and production database management.
ruvnet/ruflo
Agent skill for performance-optimizer - invoke with $agent-performance-optimizer
dosco/aithy
This skill helps an LLM generate correct AxAgent observability code using @ax-llm/ax.
dosco/aithy
This skill helps an LLM generate correct audio code with @ax-llm/ax.
dosco/aithy
This skill helps an LLM generate correct AxGEPA optimization code using @ax-llm/ax.
dosco/aithy
This skill helps with using the @ax-llm/ax TypeScript library for building LLM applications.
dosco/aithy
This skill helps an LLM build correct native Model Context Protocol integrations with @ax-llm/ax.
dosco/aithy
This skill helps an LLM generate correct playbook code using @ax-llm/ax.
This skill helps an LLM generate correct AxAgent tuning and evaluation code using @ax-llm/ax. Ax Agent Optimize is an agent skill from dosco/aithy. This skill helps an LLM generate correct AxAgent tuning and evaluation code using @ax-llm/ax.
Ax Agent Optimize fits situations like: the user asks about agent.optimize(...); optimization targets; saved optimizedProgram artifacts; agent optimization guidance.
Run `npx skills add dosco/aithy --skill ax-agent-optimize -a claude-code`. Or copy the skill folder (.claude/skills/ax-agent-optimize in dosco/aithy) into .claude/skills/ax-agent-optimize in your project. Claude Code loads it when a task matches its description.
Run `npx skills add dosco/aithy --skill ax-agent-optimize -a codex`. Or copy the skill folder (.claude/skills/ax-agent-optimize in dosco/aithy) into .agents/skills/ax-agent-optimize in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dosco/aithy --skill ax-agent-optimize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ax-agent-optimize, .gemini/skills/ax-agent-optimize, .github/skills/ax-agent-optimize and .opencode/skills/ax-agent-optimize in your project.
SKILL.md names no scripts, command-line tools or credentials: Ax Agent Optimize is instructions for the agent only.
SKILL.md names 1 domain. As links in the text: raw.githubusercontent.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ax Agent Optimize is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ax Agent Optimize: Caveman Optimization Evaluator (JuliusBrussee/caveman, 110k stars), SQL Optimization (github/awesome-copilot, 40k stars), Database Optimizer (davila7/claude-code-templates, 32k stars) and Arize Evaluator (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
dosco (a GitHub user) maintains it in dosco/aithy, which has 107 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on August 31, 2026.
Source: dosco/aithy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.