Agent skill

Agent Wiki Compare Outcomes

by AgentToolkit in AgentToolkit/altk-evolve

Contrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments.

Apache-2.0Auto-check passedAgent Workflows

Install Agent Wiki Compare Outcomes

skills CLI
$ npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AgentToolkit/altk-evolve agent-wiki-compare-outcomes --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AgentToolkit/altk-evolve.git skills-src && mkdir -p .claude/skills && cp -r skills-src/explorations/agent-wiki/skills/agent-wiki-compare-outcomes .claude/skills/agent-wiki-compare-outcomes && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-wiki-compare-outcomes
GitHub stars
122
Token cost
~1.7k tokens
SKILL.md length
612 words
Files
2 (incl. scripts)
Skills in repo
21
Repo updated
First seen
Licence
Apache-2.0

At a glance

Contrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments.

  • Works in 3 steps: Build an Evidence Pack → Inspect Candidate Rules → Promote Carefully
  • Learning rules from multiple successful and failed runs of the same task
  • SKILL.md covers Overview, Workflow and Guardrails
  • Runs Python scripts from its folder; calls uv

What it does

Used after the summarize, extract and synthesize passes, this step looks at several normalized trajectories of the same or similar task and sets those that succeeded against those that failed. A bundled Python script, compare_outcomes.py, groups runs by task id, falling back to the normalized task request, and builds an evidence pack: request text, stored outcomes and failure snippets, observed tool and API calls, and tool descriptions shown in the transcript.

Outcomes can come from stored labels or from an LLM judging the transcript, selected with a --judge-outcomes flag set to never, missing or always. The always setting is preferred when stored labels come from a benchmark-specific evaluator. A rule is written only when the trajectories contain evidence for it, so guidelines are never filled in from prior knowledge of the benchmark or application.

When your agent uses it

  • Learning rules from multiple successful and failed runs of the same task
  • Turning benchmark trajectories into evidence-backed agent guidelines
  • Judging run outcomes with an LLM when evaluator labels are missing or benchmark-specific

Example prompts

  • “Compare the successful and failed runs in ./trajectories/normalized and propose guidelines.”
  • “Run compare_outcomes.py with --judge-outcomes always on both experiment folders and summarize the contrasts.”

Requirements

  • uv and Python to run compare_outcomes.py
  • Normalized trajectory JSON files

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Build an Evidence Pack
  2. Inspect Candidate Rules
  3. Promote Carefully

What it can do on your machine

Read from SKILL.md and the folder at commit 9e5bb56. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Wiki Compare Outcomes loads about 1.7k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 612 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from AgentToolkit/altk-evolve at commit 9e5bb56, republished under its Apache-2.0 licence (© AgentToolkit). 612 words, ~1,654 tokens.

Download SKILL.mdSave it as .claude/skills/agent-wiki-compare-outcomes/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
agent-wiki-compare-outcomes
description
Compare successful and failed normalized agent trajectories to derive evidence-backed agent-wiki guidelines. Use when Codex has multiple runs for the same or similar task, evaluator outcomes, failed/successful variants, benchmark trajectories, or wants to learn rules from contrasts rather than from one trajectory alone.

Agent Wiki — Compare Outcomes

Overview

Use this pass after summarize/extract/synthesize when there are multiple trajectories that can be judged as successful or failed. It can judge outcomes with an LLM from the normalized transcript, so it does not need to depend on benchmark-specific success/failure labels. It derives contrastive guidelines: rules that are supported by a failed path, a successful path, and concrete evidence from task wording, tool/API documentation, tool/API calls, transcript evidence, optional failure snippets, and optionally an LLM success/failure judgment.

This pass exists to avoid hand-authored domain knowledge. Do not write a rule just because you know the benchmark or application. Write a rule only when the input trajectories contain the evidence.

Workflow

Step 1: Build an Evidence Pack

Run the bundled script over normalized trajectory JSON files:

bash
uv run python explorations/agent-wiki/skills/agent-wiki-compare-outcomes/scripts/compare_outcomes.py \
  --input <normalized-dir-or-json> \
  --out-json <analysis.json> \
  --out-md <analysis.md> \
  --judge-outcomes always

Pass --input multiple times to compare several experiment arms.

The script groups traces by metadata.task_id when present; otherwise it uses a normalized task request. For each group it compares successful and failed runs, then extracts:

  • task request text;
  • stored outcome and failure snippets when present;
  • LLM-judged outcome when --judge-outcomes missing or --judge-outcomes always is set;
  • observed tool/API calls from stats.top_tools, code snippets, and source api_calls.jsonl when available;
  • tool/API descriptions shown in the trajectory transcript.

Judging modes:

  • --judge-outcomes never: use only stored outcome.success.
  • --judge-outcomes missing: judge only traces without stored outcomes.
  • --judge-outcomes always: ignore stored success labels and use the LLM judgment for all traces.

Prefer --judge-outcomes always when the available stored labels come from a benchmark evaluator or another dataset-specific schema. Use stored outcomes only when they are trusted, dataset-neutral annotations you are comfortable using as ground truth.

Use --judge-include-failures when generic failure reports or evaluator snippets are available and you want the LLM to interpret them. This does not require benchmark-specific code; the snippets are passed as opaque evidence. Without failure snippets or ground truth, an LLM can still identify obvious tool errors, step-limit failures, missing finalization, or apparent success, but it may not detect silent semantic mismatches.

Show full SKILL.md (284 more words)Show less
Step 2: Inspect Candidate Rules

Read the generated Markdown. A candidate is promotable only if it has:

  • at least one failed trajectory and one successful trajectory in the same group;
  • a task-action tool/API or workflow difference between them, not just authentication, documentation lookup, or finalization calls;
  • a comparison between plausible alternatives in the same tool namespace, unless the transcript evidence clearly supports a cross-namespace workflow rule;
  • evidence that the successful tool/API is more semantically aligned with the current task wording, or that the transcript/failure evidence names the failed side effect;
  • source trajectory IDs for both sides.

If the evidence is incomplete, keep it as a hypothesis. Hypotheses are useful for evaluation notes but should not be promoted into future-agent instructions.

Step 3: Promote Carefully

When a candidate is strong, render it as a guideline with provenance:

json
{
  "entities": [
    {
      "type": "guideline",
      "title": "Choose record source from task wording",
      "content": "Apply this rule only when the live choice is between the observed successful and failed APIs, or between APIs with the same documented meanings. Prefer the successful source when the request matches its observed documentation. Do not apply this rule when the request explicitly uses failed-side terms; inspect the failed-side source instead. Do not generalize this rule to other record families or unrelated APIs unless a separate contrast includes those APIs.",
      "rationale": "In the contrasted trajectories, failed runs used a feed endpoint for a task about the user's own transactions, while the successful run used the documented account-owned transaction endpoint.",
      "trigger": "Use only when choosing between the observed successful and failed APIs and the task wording aligns with the successful-side documentation; skip when the task explicitly mentions failed-side terms or asks about a different record family.",
      "session_id": "<comparison-id>",
      "agent": "agent-wiki-compare-outcomes",
      "tags": ["contrastive", "tool-selection", "data-source-routing"],
      "normalized_path": "<analysis.json>"
    }
  ]
}

Pipe through the normal helper:

bash
cat /tmp/contrastive-guideline.json | uv run python explorations/agent-wiki/skills/scripts/build_agent_wiki.py --wiki-root <wiki-root> render-guidelines
uv run python explorations/agent-wiki/skills/scripts/build_agent_wiki.py --wiki-root <wiki-root> catalog

Guardrails

  • Do not derive rules from outcome labels or private evaluator data alone. Outcome labels and LLM judgments can identify which side failed, but the proposed future behavior must come from trajectory-visible task wording, observed calls, or observed documentation.
  • Do not invent tool/API names. A concrete name must appear in a call or in retrieved documentation.
  • Prefer generic rule wording first, with tool-specific examples under evidence. The wiki can specialize only where the evidence supports it.
  • Keep triggers narrow. Name the observed successful and failed API pair, add the successful-side positive terms, and add explicit counter-scope for failed-side terms and unrelated record families.
  • Record counterexamples: if a failed and successful run used the same tool, this pass did not identify a source-selection rule.
  • Keep confidence explicit. High confidence requires at least one success, one failure, and a clear successful-only vs failed-only behavior difference.

© AgentToolkit, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in explorations/agent-wiki/skills/agent-wiki-compare-outcomes of AgentToolkit/altk-evolve.

  • SKILL.md
  • scripts/compare_outcomes.py

Open the folder on GitHubat commit 9e5bb56

Compare with similar skills

Agent Wiki Compare Outcomes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Wiki Compare Outcomes compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Wiki Compare Outcomes this skillAgentToolkit/altk-evolve122—~1.7kAutomated safety check: PassApache-2.0
MCP Server Builderanthropics/skills180k62 repos~2.3kAutomated safety check: PassApache-2.0
Mem0 CLI Memory Commandsmem0ai/mem067k—~1.4kAutomated safety check: NotesApache-2.0
MemPalace Setup and OperationMemPalace/mempalace59k—~2.2kAutomated safety check: PassMIT
SkillOpt-Sleep Self-Improvement Cyclemicrosoft/SkillOpt18k—~2.3kAutomated safety check: PassMIT
Agent Memorytigerless-labs/agent-memory2.4k—~1.2kAutomated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 62 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Adds, searches, lists, updates and deletes memories on the Mem0 platform from the terminal with the mem0 command, including a JSON mode built for agents.

    67k GitHub stars~1.4k tokensUpdated yesterday
    Agent WorkflowsAuto-check: notes
  • Installs and configures MemPalace as a private local palace, a shared-brain hub or a client of an existing hub, including MCP registration and version-correct initialization.

    59k GitHub stars~2.2k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Official

    Runs a nightly or on-demand sleep cycle for a local Codex agent: review past sessions, replay recurring tasks and stage validated skill and memory edits for adoption.

    18k GitHub stars~2.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Agent Memory

    tigerless-labs/agent-memory

    Read and write the shared long-term memory store. An agent skill from tigerless-labs/agent-memory.

    2.4k GitHub stars~1.2k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Mindmemos CLI

    mindscale-noah/MindMemOS

    Give an AI agent persistent, cross-session long-term memory through MindMemOS.

    1k GitHub stars~3.4k tokensUpdated 8 days ago
    Agent WorkflowsAuto-check passed

More from AgentToolkit/altk-evolve

All 21 skills in this repo
  • Trajectory Skill Synthesizer

    AgentToolkit/altk-evolve

    Turns a saved agent trajectory into a reusable skill with a SKILL.md and supporting scripts, so later sessions can call the workflow instead of rediscovering it.

    122 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Evolve Memory Mirror

    AgentToolkit/altk-evolve

    Copies a memory the agent just saved into the shared evolve store with a retrieval trigger, so it can be shared across a team and audited like other entities.

    122 GitHub stars~922 tokensUpdated today
    Auto-check passed
  • Agent Wiki Consult

    AgentToolkit/altk-evolve

    Has the agent read a project wiki's AGENTS.md and pull the guidelines that fit the task it is about to start, instead of loading the wiki at session start.

    122 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Wiki Guideline Extractor

    AgentToolkit/altk-evolve

    Reads a normalized Claude Code trajectory JSON and turns its errors and solutions into reusable guideline pages in wiki-twobatch/guidelines.

    122 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Agent Wiki Ingest

    AgentToolkit/altk-evolve

    Ingest one or more agent trajectories (raw bob/claude traces or normalized JSON) into an agent-wiki end-to-end — convert, summarize, extract guidelines, synthesize skills, optionally compare…

    122 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Agent Wiki Summarize

    AgentToolkit/altk-evolve

    Read a normalized Claude Code trajectory JSON and write an episodic summary page to wiki-twobatch/summaries/.

    122 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Agent Wiki Compare Outcomes

What does Agent Wiki Compare Outcomes do?

Contrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments. Used after the summarize, extract and synthesize passes, this step looks at several normalized trajectories of the same or similar task and sets those that succeeded against those that failed.py, groups runs by task id, falling back to the normalized task request, and builds an evidence pack: request text, stored outcomes and failure snippets, observed tool and API calls, and tool descriptions shown in the transcript.

When should I use Agent Wiki Compare Outcomes?

Agent Wiki Compare Outcomes fits situations like: learning rules from multiple successful and failed runs of the same task; turning benchmark trajectories into evidence-backed agent guidelines; judging run outcomes with an LLM when evaluator labels are missing or benchmark-specific.

How do I install Agent Wiki Compare Outcomes in Claude Code?

Run `npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a claude-code`. Or copy the skill folder (explorations/agent-wiki/skills/agent-wiki-compare-outcomes in AgentToolkit/altk-evolve) into .claude/skills/agent-wiki-compare-outcomes in your project. Claude Code loads it when a task matches its description.

How do I install Agent Wiki Compare Outcomes in Codex?

Run `npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a codex`. Or copy the skill folder (explorations/agent-wiki/skills/agent-wiki-compare-outcomes in AgentToolkit/altk-evolve) into .agents/skills/agent-wiki-compare-outcomes in your project. Codex loads it when a task matches its description.

Can I use Agent Wiki Compare Outcomes in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-wiki-compare-outcomes, .gemini/skills/agent-wiki-compare-outcomes, .github/skills/agent-wiki-compare-outcomes and .opencode/skills/agent-wiki-compare-outcomes in your project.

What does Agent Wiki Compare Outcomes need to run?

Going by SKILL.md and its folder, Agent Wiki Compare Outcomes needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: uv and Python to run compare_outcomes.py; Normalized trajectory JSON files.

Does Agent Wiki Compare Outcomes access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Agent Wiki Compare Outcomes safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agent Wiki Compare Outcomes use?

Agent Wiki Compare Outcomes is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Wiki Compare Outcomes use?

About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Wiki Compare Outcomes?

Skills that share tags, products or a category with Agent Wiki Compare Outcomes: MCP Server Builder (anthropics/skills, 180k stars), Mem0 CLI Memory Commands (mem0ai/mem0, 67k stars), MemPalace Setup and Operation (MemPalace/mempalace, 59k stars) and SkillOpt-Sleep Self-Improvement Cycle (microsoft/SkillOpt, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Wiki Compare Outcomes?

AgentToolkit (a GitHub organization) maintains it in AgentToolkit/altk-evolve, which has 122 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 7, 2026.

Source: AgentToolkit/altk-evolve on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.