Agent skill

Ln 45 Benchmark Comparator

by levnikolaevich in levnikolaevich/claude-code-skills

Compares tools or implementations through controlled benchmarks and independent correctness checks.

MITAuto-check passed

Install Ln 45 Benchmark Comparator

skills CLI
$ npx skills add levnikolaevich/claude-code-skills --skill ln-45-benchmark-comparator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install levnikolaevich/claude-code-skills ln-45-benchmark-comparator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/levnikolaevich/claude-code-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/implementation-suite/skills/ln-45-benchmark-comparator .claude/skills/ln-45-benchmark-comparator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ln-45-benchmark-comparator
GitHub stars
574
Token cost
~3.6k tokens
SKILL.md length
1,843 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

Compares tools or implementations through controlled benchmarks and independent correctness checks.

  • Works in 5 steps: Define the Decision and Experiment → Build a Symmetric Harness → Execute and Capture Evidence → …
  • SKILL.md covers Tool Routing, Evidence Rules, Checklist and Self-Check, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Ln 45 Benchmark Comparator is an agent skill from levnikolaevich/claude-code-skills. Compares tools or implementations through controlled benchmarks and independent correctness checks.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Help your AI agent finish the job: solve the right problem, keep changes focused, and show what was verified. For Claude Code and Codex. The licence is MIT.

Example prompts

  • “Use the ln-45-benchmark-comparator skill to compare tools or implementations through controlled benchmarks and independent correctness checks”
  • “/ln-45-benchmark-comparator”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Define the Decision and Experiment
  2. Build a Symmetric Harness
  3. Execute and Capture Evidence
  4. Analyze Validity and Results
  5. Decide, Preserve, and Clean Up

What it can do on your machine

Read from SKILL.md and the folder at commit 0ce8796. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ln 45 Benchmark Comparator loads about 3.6k tokens when it runs. Until then it costs about 32 tokens; SKILL.md has 1,843 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from levnikolaevich/claude-code-skills at commit 0ce8796, republished under its MIT licence (© levnikolaevich). 1,843 words, ~3,630 tokens.

Download SKILL.mdSave it as .claude/skills/ln-45-benchmark-comparator/SKILL.md (or your agent's skills folder).
name
ln-45-benchmark-comparator
description
Compares tools or implementations through controlled benchmarks and independent correctness checks.

Benchmark Comparator

Goal: Compare alternatives under controlled, reproducible conditions. Correctness comes before speed, and measured data must remain separate from estimates, setup cost, and interpretation.

Execution contract: The checklist defines completion. Track each item internally as PENDING, PROVEN with evidence, CLEARED with evidence its condition is absent, or UNPROVEN with a gap; reading, delegation, tool failure, a zero exit status, or a self-reported success is not proof; only the observed outcome is. Reconcile after each section. Before returning, resolve all PENDING, count only PROVEN and CLEARED, and apply verdict and approval rules to every gap. Preserve intent, scope, and existing authorization. Continue authorized work; ask only for consequential unresolved choices or required external approval. When no one can answer during the run, state the exact question and apply the skill's verdict for the remaining gap instead of waiting or guessing. Scale depth to material risk without skipping checks. Preserve dependency and safety order; otherwise choose an appropriate verification method. Accept equivalent user or repository evidence; no other skill, named artifact, or complete lifecycle is required. Preserve source requirement and decision IDs. Bind reused evidence to relevant source versions, dirty changes, configuration, and environment; invalidate only affected claims. On continuation, reconcile task, authorization, current state, and unresolved evidence. For long work, return a compact continuation record or update an already authorized artifact; read-only skills do not persist it. Distinguish artifact readiness, verified behavior, and external-action authority. Prepare authorized work before required approval. If blocked by an instruction, cite its exact source and unresolved boundary; do not invent approval gates from caution.

Tool Routing

NeedPreferred toolUse it whenFallback
Canonical workload and oracleRepository fixtures, tests, expected diffs, schemas, or independently specified outcomesDefining what success means before either candidate runsCreate the smallest deterministic fixture that represents the decision
IsolationClean Git worktrees, temporary directories, controlled environment, fixed seeds, and resettable cachesPreventing one candidate or run from contaminating anotherSequential clean-room setup with verified cleanup
ExecutionThe same shell runner and wrapper for every candidateCapturing commands, exit status, stdout, stderr, timing, and artifacts consistentlyManual execution with an explicit reproducibility limitation
Activation proofLogs, traces, command records, process metadata, or candidate-specific artifactsVerifying the intended alternative actually ran and did not fall backTreat the run as invalid when activation cannot be proven
Correctness gradingTests, output parser, diff, schema validation, or independent oracleEvery scenario before cost comparisonManual blind grading against written expectations
Performance and costMonotonic timer, resource metrics, token or usage telemetry, tool-call logs, and failure countsMetrics are observable through the same method for all candidatesLabel derived or estimated values and keep them out of measured aggregates
External semanticsOfficial documentation and specificationsCandidate configuration or claimed behavior needs current verificationPrimary-source web research; otherwise mark the claim UNVERIFIED

Do not tune the scenario after observing a preferred candidate, mix measurements from different workloads, or present internal estimates as externally measured facts. Benchmarking may create temporary worktrees and artifacts but must not change the source baseline or unapproved external state.

Evidence Rules

  • Hold all non-tested variables constant or record and analyze the confounder.
  • Correctness failure cannot be compensated by better speed, token use, or cost unless the decision explicitly allows degraded correctness.
  • Use repeated runs and report raw values, center, spread, failures, and outliers; never headline the best run.
  • Keep setup or indexing cost, steady-state cost, maintenance burden, and runtime cost separate.
  • Report measured, derived, estimated, and qualitative evidence in distinct fields.

Checklist

1. Define the Decision and Experiment
  • State the decision the benchmark must support, the candidates, intended users, representative workloads, and explicit non-goals.
  • Include ordinary cases where a simpler or built-in candidate could reasonably win as well as cases exercising each candidate's claimed advantage; do not construct a feature demo for one side.
  • Define scenario inputs, expected outcomes, correctness criteria, failure conditions, and an oracle independent of candidate self-report.
  • Define primary and secondary metrics, units, measurement point, acceptance threshold, allowed tradeoffs, and tie or inconclusive rules.
  • Define the treatment variable, then hold other relevant variables fixed: revision, model, prompts, permissions, runtime, hardware, data, caches, and network. A variable being compared cannot also be declared fixed.
  • Predeclare repetitions using existing variability evidence or a bounded pilot, or a bounded sequential rule with fixed error tolerance and maximum runs. Return INCONCLUSIVE if the budget cannot resolve the meaningful effect.
  • Freeze and hash the scenario text, fixtures, expectations, runner, parser, and decision rule before the first candidate result is inspected.
  • Read repository instructions and inspect Git state before creating worktrees, temporary data, or runners.
  • Start a run-owned resource ledger with every created absolute path, worktree, process ID, cache, account, dataset, report, and temporary artifact; never register pre-existing resources or credentials as cleanup targets.
  • Require side-effect-free or idempotent scenarios and disposable accounts or datasets; define external-write authorization, cost and rate budgets, cleanup, and rollback evidence before execution.
  • Return BLOCKED if candidates do not solve the same task, correctness cannot be independently graded, external effects cannot be isolated, or the decision rule is being chosen after results.
2. Build a Symmetric Harness
  • Use the same runner, timeout, logging, environment construction, and artifact collection for every candidate.
  • Create clean worktrees or equivalent isolated copies from the same commit and verify identical starting state.
  • Inventory global and user-level instructions, hooks, plugins, settings, credentials, caches, and environment variables that can leak candidate behavior across supposedly isolated arms; disable, equalize, or record each one.
  • Control seeds, clock, locale, concurrency, network access, dependency versions, cache state, warmup, and scenario order where they can influence results.
  • Record exact candidate configuration, feature flags, prompts, command lines, permissions, and versions.
  • Define a symmetric tuning policy and budget—default configuration, equally tuned configuration, or both—so one candidate is not optimized after seeing the other's result.
  • Add an activation check that proves each candidate was used and identifies silent fallback, partial activation, or mixed execution.
  • Validate the output parser, diff rules, test oracle, and metric collector on known pass and fail fixtures before benchmarking.
  • Separate one-time setup, indexing, compilation, or download cost from steady-state execution and amortized cost.
  • Define cleanup and failure recovery from the resource ledger so a crashed or timed-out run cannot contaminate later runs.
Show full SKILL.md (811 more words)Show less
3. Execute and Capture Evidence
  • Run candidates in a balanced or randomized order that avoids systematic warm-cache or temporal advantage.
  • Capture start and end state, command, exit status, timing, resource metrics, logs, outputs, diffs, tests, and candidate-specific artifacts for every run.
  • Verify activation before grading; mark unproven or fallback runs invalid rather than assigning them to the intended candidate.
  • Grade correctness against the predefined oracle before examining performance and cost metrics.
  • Grade task completeness separately from correctness and efficiency: verify every required outcome and prohibited side effect, rather than treating a smaller diff, lower token count, or successful subset as completion.
  • Blind manual or qualitative graders to candidate identity and randomize presentation order; record disagreements instead of resolving them toward a preferred candidate.
  • Record timeout, crash, malformed output, partial completion, tool error, and environmental failure as distinct failure classes.
  • Follow the predeclared repetition or sequential stopping rule; count and preserve all attempts, including failed and invalid runs. Do not keep retrying until the desired number of successes appears.
  • Pause when environmental drift, rate limits, external outages, background load, or runner defects make additional runs incomparable.
  • After a harness fix, invalidate affected comparisons and rerun both candidates for those scenarios under the corrected harness; retain unaffected evidence and label prior invalid results.
  • Treat setup, activation, parser, and environmental failures separately from task incorrectness, then state whether setup reliability is part of the actual product decision.
4. Analyze Validity and Results
  • Exclude only runs that meet a predefined invalidation rule and record the reason, evidence, and whether exclusion changes the conclusion.
  • Report per-scenario correctness, failures, latency, resource use, tokens or usage, tool calls, and other costs before aggregating.
  • Use median, percentile, confidence interval, or another statistic appropriate to the sample and distribution; show spread and sample size.
  • Keep metrics with different units or workloads separate and avoid a single composite score unless its weighting was defined before execution.
  • Check whether differences exceed measurement noise and whether one scenario dominates the aggregate.
  • Analyze setup cost, steady-state cost, maintenance complexity, portability, failure behavior, and operational burden separately from runtime metrics.
  • Label synthetic fixtures, estimated tokens, character-based proxies, modeled cost, and manual judgments so they cannot be mistaken for observed telemetry.
  • Request an independent blind review when qualitative output quality materially affects the decision and automated correctness is insufficient.
5. Decide, Preserve, and Clean Up
  • Identify which decision the comparison can support, its valid workload/environment range, and conditions requiring another experiment; do not extrapolate outside observed evidence.
  • Use WIN only when the candidate satisfies correctness and the predefined decision rule with sufficient valid evidence.
  • Use TIE only when evidence supports the predefined negligible-difference margin or balanced tradeoff; failure to detect a difference with insufficient evidence is INCONCLUSIVE.
  • Use INCONCLUSIVE when sample size, activation, oracle, environmental control, or conflicting scenarios prevent a reliable choice.
  • Preserve reproducible commands, configuration, scenario definitions, expectations, raw results, normalized results, and analysis needed for independent verification.
  • Remove only run-owned ledger entries: verify absolute paths remain inside approved temporary roots, stop exact recorded process IDs, preserve dirty or pre-existing worktrees, never delete credentials, and verify source and external baseline state.
  • Report invalid runs, exclusions, confounders, sensitivity to assumptions, and how the conclusion could be falsified.
  • Report residual decision risks that remain after the comparison, including unsupported workloads, unmeasured costs, unstable environments, and assumptions that could reverse the verdict.
  • Verify that decision guidance follows scenario-level evidence and the frozen rule, with cleanup and limitations accounted for.

Self-Check

  • Reconcile before returning. Check item-level evidence, requirement coverage, contradictions, scope, verdict, and applicable cleanup. Correct the report or authorized artifacts. Reuse valid evidence; do not automatically rescan the repository or rerun successful commands. Repeat checks only for relevant changes, failures, or unresolved evidence. Disclose remaining gaps.

Output Contract

Report in the user's language, in this order; label all five fields and state each fact once. Use controlled plain language: one fact per sentence, usually under 20 words, active voice, and one term per concept, with no synonyms for verdicts, IDs, or states. Small results may use one line per field; omit empty tables and do not copy linked artifacts:

  1. Result: The exact skill-specific verdict token first, then the supported outcome.
  2. Scope: Reviewed/changed scope, exclusions, baseline, and material assumptions.
  3. Evidence: Skill-specific fields below; distinguish facts, inferences, and unverified claims. Link artifacts; use tables when useful.
  4. Verification: Checks/results, unavailable evidence, and applicable cleanup/external state.
  5. Completion: Checklist: X/Y complete; Incomplete: None or each UNPROVEN item's reason, outcome impact, and exact next action; residual risks and required decisions.

Skill-specific evidence: Frozen decision, candidates, scenarios, independent oracle, fixed variables, configuration, metrics, repetitions, exclusions, and decision rule. Prove activation and harness validity; report invalid runs, confounders, exclusions, and cleanup. Compare scenario-level completeness, correctness, failures, primary metric, spread/sample size, and other costs before aggregating. Preserve setup/maintenance tradeoffs, sensitivity, falsification conditions, raw results, configurations, scenario artifacts, and hashes when available.

© levnikolaevich, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/implementation-suite/skills/ln-45-benchmark-comparator of levnikolaevich/claude-code-skills.

Open the folder on GitHubat commit 0ce8796

Compare with similar skills

Ln 45 Benchmark Comparator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ln 45 Benchmark Comparator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ln 45 Benchmark Comparator this skilllevnikolaevich/claude-code-skills574—~3.6kAutomated safety check: PassMIT
Implementing Endpoint Dlp Controlsmukul975/Anthropic-Cybersecurity-Skills34k—~1.5kAutomated safety check: PassApache-2.0
Implementing Network Access Controlmukul975/Anthropic-Cybersecurity-Skills34k—~3.8kAutomated safety check: NotesApache-2.0
Implementing Usb Device Control Policymukul975/Anthropic-Cybersecurity-Skills34k—~1.4kAutomated safety check: PassApache-2.0
Implementing Pci Dss Compliance Controlsmukul975/Anthropic-Cybersecurity-Skills34k—~1.6kAutomated safety check: PassApache-2.0
Implementing Nerc Cip Compliance Controlsmukul975/Anthropic-Cybersecurity-Skills34k—~4.3kAutomated safety check: PassApache-2.0

Similar skills

  • Implementing Endpoint Dlp Controls

    mukul975/Anthropic-Cybersecurity-Skills

    Implements endpoint Data Loss Prevention (DLP) controls to detect and prevent sensitive data exfiltration through email, USB, cloud storage, and printing.

    34k GitHub stars~1.5k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Implementing Network Access Control

    mukul975/Anthropic-Cybersecurity-Skills

    Implements 802.1X port-based network access control using RADIUS authentication, PacketFence NAC, and switch configuration to enforce identity-based access policies, posture assessment, and…

    34k GitHub stars~3.8k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check: notes
  • Implementing Usb Device Control Policy

    mukul975/Anthropic-Cybersecurity-Skills

    Implements USB device control policies to restrict unauthorized removable media access on endpoints, preventing data exfiltration and malware introduction via USB devices.

    34k GitHub stars~1.4k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Implementing Pci Dss Compliance Controls

    mukul975/Anthropic-Cybersecurity-Skills

    Implements PCI DSS 4.0.1's 12 requirements across 6 control objectives for organizations that store, process, or transmit cardholder data, including the customized validation approach, enhanced…

    34k GitHub stars~1.6k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Implementing Nerc Cip Compliance Controls

    mukul975/Anthropic-Cybersecurity-Skills

    Implements NERC CIP controls for Bulk Electric System (BES) cyber systems: asset categorization (CIP-002), electronic security perimeters (CIP-005), system security management (CIP-007)…

    34k GitHub stars~4.3k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Benchmark

    affaan-m/ECC

    Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…

    276k GitHub starsUsed in 3 repos~654 tokens
    Frontend & DesignAuto-check passed

More from levnikolaevich/claude-code-skills

All 31 skills in this repo
  • Ln 53 Documentation Auditor

    levnikolaevich/claude-code-skills

    Audits documentation and comments for trust, coverage, consistency and freshness; read-only.

    574 GitHub stars~3.8k tokensUpdated 5 days ago
    Auto-check passed
  • Ln 81 Skill Reviewer

    levnikolaevich/claude-code-skills

    Reviews skill instructions, trigger boundaries and distribution contracts; not product code.

    574 GitHub stars~3.5k tokensUpdated 5 days ago
    Auto-check passed
  • Ln 11 Opportunity Evaluator

    levnikolaevich/claude-code-skills

    Evaluates new product opportunities through demand, channels and economics before committing to build.

    574 GitHub stars~3k tokensUpdated 5 days ago
    Auto-check passed
  • Ln 12 Product Requirements Builder

    levnikolaevich/claude-code-skills

    Defines product requirements, business rules and acceptance criteria for a committed intent; edits product docs only.

    574 GitHub stars~1.9k tokensUpdated 5 days ago
    Auto-check passed
  • Ln 13 Interaction Design Builder

    levnikolaevich/claude-code-skills

    Designs user flows, interaction states and mockups for a defined product scope; does not implement UI code.

    574 GitHub stars~1.8k tokensUpdated 5 days ago
    Auto-check passed
  • Ln 21 System Design Baseline Builder

    levnikolaevich/claude-code-skills

    Defines measurable architecture drivers and constraints before system design; edits architecture docs only.

    574 GitHub stars~2.5k tokensUpdated 5 days ago
    Auto-check passed

Questions about Ln 45 Benchmark Comparator

What does Ln 45 Benchmark Comparator do?

Compares tools or implementations through controlled benchmarks and independent correctness checks. Ln 45 Benchmark Comparator is an agent skill from levnikolaevich/claude-code-skills. Compares tools or implementations through controlled benchmarks and independent correctness checks.

How do I install Ln 45 Benchmark Comparator in Claude Code?

Run `npx skills add levnikolaevich/claude-code-skills --skill ln-45-benchmark-comparator -a claude-code`. Or copy the skill folder (plugins/implementation-suite/skills/ln-45-benchmark-comparator in levnikolaevich/claude-code-skills) into .claude/skills/ln-45-benchmark-comparator in your project. Claude Code loads it when a task matches its description.

How do I install Ln 45 Benchmark Comparator in Codex?

Run `npx skills add levnikolaevich/claude-code-skills --skill ln-45-benchmark-comparator -a codex`. Or copy the skill folder (plugins/implementation-suite/skills/ln-45-benchmark-comparator in levnikolaevich/claude-code-skills) into .agents/skills/ln-45-benchmark-comparator in your project. Codex loads it when a task matches its description.

Can I use Ln 45 Benchmark Comparator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add levnikolaevich/claude-code-skills --skill ln-45-benchmark-comparator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ln-45-benchmark-comparator, .gemini/skills/ln-45-benchmark-comparator, .github/skills/ln-45-benchmark-comparator and .opencode/skills/ln-45-benchmark-comparator in your project.

What does Ln 45 Benchmark Comparator need to run?

SKILL.md names no scripts, command-line tools or credentials: Ln 45 Benchmark Comparator is instructions for the agent only.

Does Ln 45 Benchmark Comparator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ln 45 Benchmark Comparator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ln 45 Benchmark Comparator use?

Ln 45 Benchmark Comparator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ln 45 Benchmark Comparator use?

About 3.6k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ln 45 Benchmark Comparator?

Skills that share tags, products or a category with Ln 45 Benchmark Comparator: Implementing Endpoint Dlp Controls (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Implementing Network Access Control (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Implementing Usb Device Control Policy (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Implementing Pci Dss Compliance Controls (mukul975/Anthropic-Cybersecurity-Skills, 34k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ln 45 Benchmark Comparator?

levnikolaevich (a GitHub user) maintains it in levnikolaevich/claude-code-skills, which has 574 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 5, 2026.

Source: levnikolaevich/claude-code-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.