Agent skill

Context Compare

by ai-analyst-lab in ai-analyst-lab/ai-analyst

Advanced: runs the same question under two configurations and diffs the results.

MITAuto-check passed

Install Context Compare

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill context-compare -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst context-compare --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/context-compare .claude/skills/context-compare && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
context-compare
GitHub stars
304
Token cost
~1.8k tokens
SKILL.md length
956 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Advanced: runs the same question under two configurations and diffs the results.

  • Works in 5 steps: back up the metric dictionary → baseline arm (without the context) → with-context arm (definition staged) → …
  • /context-compare
  • SKILL.md covers Purpose, Invocation, How to run it and Notes
  • Calls python3

What it does

Context Compare is an agent skill from ai-analyst-lab/ai-analyst. Advanced: runs the same question under two configurations and diffs the results. Ask one analytics question with a piece of context and without it, and measure what changed. Trigger on "/context-compare", "run it with and without <the definition/context", "does adding <X change the answer", "is this context worth it".

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

When your agent uses it

  • /context-compare
  • Run it with and without <the definition/context
  • Does adding <X change the answer
  • Is this context worth it

Example prompts

  • “/context-compare”
  • “run it with and without <the definition/context”
  • “does adding <X change the answer”
  • “/context-compare”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. back up the metric dictionary
  2. baseline arm (without the context)
  3. with-context arm (definition staged)
  4. compute the delta and report
  5. restore and report

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Context Compare loads about 1.8k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 956 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 956 words, ~1,814 tokens.

Download SKILL.mdSave it as .claude/skills/context-compare/SKILL.md (or your agent's skills folder).
name
context-compare
description
Advanced: runs the same question under two configurations and diffs the results. Ask one analytics question with a piece of context and without it, and measure what changed. Trigger on "/context-compare", "run it with and without <the definition/context>", "does adding <X> change the answer", "is this context worth it".

Skill: Context compare (with and without)

Purpose

Measure what a piece of context is doing, instead of asserting it. Run the same question two ways, once without the context (for example, no metric definition) and once with it, and report the delta: did the spread collapse, did the runs start citing the definition, did the verdict go from drifts to stable. The setup whose presence collapses the drift is the context that moves the answer. Convergence is stability, not correctness.

Invocation

/context-compare "<the question>" --with <the definition> [N] Default N = 5. The context under test is a meaning-only definition, usually a metric definition: the user states it in conversation, points at a definition YAML file, or names an entry already in the metric dictionary. The baseline is the analyst with that definition absent from the dictionary.

Example: /context-compare "What's our retention rate?" --with "Retention Rate (30d): share of accounts at least 30 days old that were active in the trailing 30 days"

How to run it

This skill is glue plus bookkeeping, and the whole procedure is manual: you stage the definition into the active dataset's metric dictionary, run the reliability procedure once per arm, compute the delta between the two arms, and restore the dictionary. The run step is the reliability skill (a sibling of this one); the per-arm statistics come from that skill's bundled script. The user never types a command; you do each step. This skill ships no scripts of its own.

Script path. The stats script is bundled with the reliability skill: helpers/stats/reliability_stats.py (the same module the reliability skill uses). If the import fails (some sandboxed environments): read the script file from the reliability skill, write a copy into a scripts/ folder inside the working folder, and run it from there. The script is self-contained. In every case run it from the working-folder root, so its audit-log append lands in .knowledge/reliability/log.jsonl.

Step 0 - back up the metric dictionary

Everything below edits .knowledge/datasets/{active}/metrics/. Before touching anything, copy the whole metrics/ directory aside, for example to .knowledge/comparisons/<question-slug>/<UTC-timestamp>/metrics-backup/. This backup is what Step 4 restores; nothing this skill stages may outlive the comparison.

Step 1 - baseline arm (without the context)

Make sure the definition under test is absent: read metrics/index.yaml, and if the metric is already defined there, remove its index entry and its {id}.yaml for this arm (the Step 0 backup keeps the original).

Then run the reliability check for this arm: launch N fresh, independent sub-sessions in parallel using the reliability skill's Step 1 brief verbatim (same fresh-context rules, each sub-session sees only the question and returns the same headline / measured / definition_source block; for a comparison they must also not read .knowledge/comparisons/ before answering). Record the N results to .knowledge/comparisons/<question-slug>/<ts>-baseline/runs.json, in the same shape the reliability skill uses: {"question": "<the question>", "runs": [{"run":1,"headline":"...","measured":"...","definition_source":"..."}, ...]}

Compute this arm's statistics deterministically with the bundled script:

python3 helpers/stats/reliability_stats.py .knowledge/comparisons/<question-slug>/<ts>-baseline

It writes stats.json and report.md into the run dir. Never estimate these numbers.

Step 2 - with-context arm (definition staged)

Stage the definition into the metric dictionary using the metric-spec registration format: write .knowledge/datasets/{active}/metrics/{id}.yaml (name, plain-English definition, formula, unit, source tables, reference SQL if given) and add the id / name / category entry to metrics/index.yaml. Meaning only: the definition says how to measure, never a result number.

Run N fresh sub-sessions again with the identical brief, into .knowledge/comparisons/<question-slug>/<ts>-with-context/runs.json, and compute that arm's statistics with the same script.

Show full SKILL.md (396 more words)Show less
Step 3 - compute the delta and report

Read the two arms' stats.json files and compute the delta inline (a few lines of Python is fine: the per-arm numbers were computed by the script; the delta is plain arithmetic over those two files, never estimated):

  • Per arm: verdict, n_distinct, agreement_rate, used_dictionary, cv, range.
  • Delta: verdict change (for example DRIFT -> STABLE), change in distinct values, agreement change, range change, citations gained (the used_dictionary difference), and moved_the_answer = true Note: both arms can be individually STABLE while the headline answer relocates between them. Report the between-arm shift as the finding whenever the arm means differ materially, regardless of the per-arm verdicts. "No change" is only correct when the answer itself did not move. when the verdict improved or the spread collapsed.

Write comparison_delta.json and comparison_report.md beside the two run dirs (.knowledge/comparisons/<question-slug>/<ts>/). The report keeps this format: a per-setup table

setupverdictdistinctagreementcited definitionCVrange

followed by the delta lines (verdict change, spread drop, citations gained, moved_the_answer).

Step 4 - restore and report

Always restore: put the Step 0 backup of metrics/ back exactly as it was, so the analyst is left as found. A staged definition never persists past the comparison; if the user wants it permanent, register it afterwards with /metric-spec.

Show the user the comparison_report.md: the per-setup table (verdict, distinct values, agreement, how many runs cited the definition) and the delta (verdict change, spread drop, citations gained, and whether the context moved the answer). Frame it:

  • Moves the answer (DRIFT -> STABLE). "Without the definition the runs drifted across N readings. With it they converged and every run cited it. Same model, same data; the context did that."
  • No change. "The spread did not move. That context was not the thing the answer needed."

Notes

  • Change the analyst's context only through Steps 0, 2, and 4 (back up, stage, restore). Never hand-edit the dictionary mid-arm, and never leave a staged definition behind when the comparison is done.
  • A metric is defined by its meaning, never a hardcoded number. No result number is ever written into a staged definition; the analyst computes it from the data each run.
  • Stability is not correctness: a wrong query is perfectly stable. Compare tells you what a piece of context changed, not whether the answer is right.
  • Works the same against local DuckDB or live Snowflake; the warehouse is the analyst's connection.

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/context-compare of ai-analyst-lab/ai-analyst.

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Context Compare next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Context Compare compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Context Compare this skillai-analyst-lab/ai-analyst304—~1.8kAutomated safety check: PassMIT
Diffsopenclaw/openclaw392k4 repos~461Automated safety check: PassMIT
Configuring Windows Defender Advanced Settingsmukul975/Anthropic-Cybersecurity-Skills34k—~3.1kAutomated safety check: PassApache-2.0
Accesslint Diffsickn33/agentic-awesome-skills47k1 repos~1.2kAutomated safety check: PassMIT
Configure Channelopenclaw/openclaw392k—~946Automated safety check: PassMIT
Motion Advancedaffaan-m/ECC275k1 repos~4.7kAutomated safety check: PassMIT

Similar skills

  • Diffs

    openclaw/openclaw

    Use the diffs tool to produce real, shareable diffs (viewer URL, file artifact, or both) instead of manual edit summaries.

    392k GitHub starsUsed in 4 repos~461 tokens
    Auto-check passed
  • Configuring Windows Defender Advanced Settings

    mukul975/Anthropic-Cybersecurity-Skills

    Configures Microsoft Defender for Endpoint (MDE) advanced protection settings including attack surface reduction rules, controlled folder access, network protection, and exploit protection.

    34k GitHub stars~3.1k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Accesslint Diff

    sickn33/agentic-awesome-skills

    Diff a live page's accessibility violations against a baseline — by default compares uncommitted changes (stash-based), or pass --branch [<name] to diff against a branch.

    47k GitHub starsUsed in 1 repo~1.2k tokens
    Frontend & DesignAuto-check passed
  • Configure Channel

    openclaw/openclaw

    Configure and prove a chat channel with non-interactive one-liners; secrets only as SecretRefs.

    392k GitHub stars~946 tokensUpdated today
    Auto-check passed
  • Motion Advanced

    affaan-m/ECC

    Advanced motion patterns for React / Next.js — drag & drop, gestures, text animations, SVG path drawing, custom hooks, imperative sequences (useAnimate), loaders, and the full API decision tree.

    275k GitHub starsUsed in 1 repo~4.7k tokens
    Auto-check passed
  • Diff Review

    nexu-io/open-design

    Render the patch-edit run's accumulated changes as a reviewable diff, surface it through a GenUI choice surface, and persist the user's accept / reject decision into the artifact manifest.

    100k GitHub stars~508 tokensUpdated today
    Frontend & DesignAuto-check passed

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 7 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 7 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 7 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed

Questions about Context Compare

What does Context Compare do?

Advanced: runs the same question under two configurations and diffs the results. Context Compare is an agent skill from ai-analyst-lab/ai-analyst. Advanced: runs the same question under two configurations and diffs the results.

When should I use Context Compare?

Context Compare fits situations like: /context-compare; run it with and without <the definition/context; does adding <X change the answer; is this context worth it.

How do I install Context Compare in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill context-compare -a claude-code`. Or copy the skill folder (.claude/skills/context-compare in ai-analyst-lab/ai-analyst) into .claude/skills/context-compare in your project. Claude Code loads it when a task matches its description.

How do I install Context Compare in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill context-compare -a codex`. Or copy the skill folder (.claude/skills/context-compare in ai-analyst-lab/ai-analyst) into .agents/skills/context-compare in your project. Codex loads it when a task matches its description.

Can I use Context Compare in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill context-compare -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/context-compare, .gemini/skills/context-compare, .github/skills/context-compare and .opencode/skills/context-compare in your project.

What does Context Compare need to run?

Going by SKILL.md and its folder, Context Compare needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Context Compare access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Context Compare safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Context Compare use?

Context Compare is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Context Compare use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Context Compare?

Skills that share tags, products or a category with Context Compare: Diffs (openclaw/openclaw, 392k stars), Configuring Windows Defender Advanced Settings (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Accesslint Diff (sickn33/agentic-awesome-skills, 47k stars) and Configure Channel (openclaw/openclaw, 392k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Context Compare?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.