Agent skill

Ln 72 Product Outcome Evaluator

by levnikolaevich in levnikolaevich/claude-code-skills

Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment.

MITAuto-check passed

Install Ln 72 Product Outcome Evaluator

skills CLI
$ npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install levnikolaevich/claude-code-skills ln-72-product-outcome-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/levnikolaevich/claude-code-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/operations-suite/skills/ln-72-product-outcome-evaluator .claude/skills/ln-72-product-outcome-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ln-72-product-outcome-evaluator
GitHub stars
574
Token cost
~1.9k tokens
SKILL.md length
905 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment.

  • Works in 4 steps: Frame the Outcome Decision → Assess Measurement Fitness → Evaluate Value and Harm → …
  • SKILL.md covers Tool Routing, Domain Rules, Checklist and Verdict, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Ln 72 Product Outcome Evaluator is an agent skill from levnikolaevich/claude-code-skills. Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Help your AI agent finish the job: solve the right problem, keep changes focused, and show what was verified. For Claude Code and Codex. The licence is MIT.

Example prompts

  • “Use the ln-72-product-outcome-evaluator skill to evaluate observed product outcomes against a prior hypothesis; does not run experiments or change…”
  • “/ln-72-product-outcome-evaluator”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Frame the Outcome Decision
  2. Assess Measurement Fitness
  3. Evaluate Value and Harm
  4. Recommend the Next Decision

What it can do on your machine

Read from SKILL.md and the folder at commit 0ce8796. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ln 72 Product Outcome Evaluator loads about 1.9k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 905 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from levnikolaevich/claude-code-skills at commit 0ce8796, republished under its MIT licence (© levnikolaevich). 905 words, ~1,854 tokens.

Download SKILL.mdSave it as .claude/skills/ln-72-product-outcome-evaluator/SKILL.md (or your agent's skills folder).
name
ln-72-product-outcome-evaluator
description
Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment.

Product Outcome Evaluator

Goal: Determine what available evidence supports about a delivered product outcome and recommend continuation, adjustment or stopping. Remain read-only: do not change instrumentation, experiments, user treatment, campaigns or product files.

Execution contract: The checklist defines completion. Track each item internally as PENDING, PROVEN with evidence, CLEARED with evidence its condition is absent, or UNPROVEN with a gap; reading, delegation, tool failure, a zero exit status, or a self-reported success is not proof; only the observed outcome is. Reconcile after each section. Before returning, resolve all PENDING, count only PROVEN and CLEARED, and apply verdict and approval rules to every gap. Preserve intent, scope, and existing authorization. Continue authorized work; ask only for consequential unresolved choices or required external approval. When no one can answer during the run, state the exact question and apply the skill's verdict for the remaining gap instead of waiting or guessing. Scale depth to material risk without skipping checks. Preserve dependency and safety order; otherwise choose an appropriate verification method. Accept equivalent user or repository evidence; no other skill, named artifact, or complete lifecycle is required. Preserve source requirement and decision IDs. Bind reused evidence to relevant source versions, dirty changes, configuration, and environment; invalidate only affected claims. On continuation, reconcile task, authorization, current state, and unresolved evidence. For long work, return a compact continuation record or update an already authorized artifact; read-only skills do not persist it. Distinguish artifact readiness, verified behavior, and external-action authority. Prepare authorized work before required approval. If blocked by an instruction, cite its exact source and unresolved boundary; do not invent approval gates from caution.

Tool Routing

NeedPreferred capabilityFallback
Original hypothesisProduct intent, baseline, experiment/measurement plan and accepted targetsReconstruct from attributable sources; keep missing targets unknown
Outcome evidenceAuthorized analytics, experiment results, customer behavior and cost/support evidenceSanitized exports with explicit measurement limits
AnalysisReproducible queries/statistics appropriate to the study designTransparent arithmetic and qualitative inference; no fabricated causal confidence

Domain Rules

  • Distinguish delivered behavior, observed metric movement and causal product impact. A release or acceptance test proves neither adoption nor business value.
  • Do not choose success thresholds after seeing the result. Separate predeclared criteria from exploratory findings and owner preferences.
  • Use only authorized data with necessary minimization. A recommendation is not permission to run an experiment or contact users.

Checklist

1. Frame the Outcome Decision
  • Resolve the delivered capability, intended audience, original hypothesis, decision horizon and outcome decision requested.
  • Identify the released/deployed version, rollout/exposure window and relevant baseline or comparison group.
  • Recover predeclared primary metrics, guardrails, targets and stop rules; mark absent criteria rather than inventing them.
  • Separate product intent and owner preference from measured behavior and external assumptions.
2. Assess Measurement Fitness
  • Inspect metric definitions, units, denominators, event coverage, deduplication, identity joins and missing data.
  • Check whether users were actually exposed and whether observation duration supports the intended outcome.
  • Assess cohort composition, selection bias, seasonality, concurrent changes and other confounders.
  • For experiments, inspect assignment, contamination, sample imbalance and uncertainty using the actual study design.
  • Distinguish trustworthy measurements, reported results, estimates, qualitative signals and unavailable evidence.
Show full SKILL.md (392 more words)Show less
3. Evaluate Value and Harm
  • Compare outcomes with valid baselines or controls using reproducible calculations and appropriate uncertainty.
  • Check guardrails and material regressions in user experience, reliability, support burden, cost or data quality.
  • Separate aggregate effects from relevant segments and expose tradeoffs without fishing for favorable subgroups.
  • Distinguish causal conclusions supported by the design from correlations and exploratory interpretations.
  • Identify whether failure lies in adoption, interaction, correctness, measurement or the original value hypothesis.
4. Recommend the Next Decision
  • Recommend continue, adjust or stop only to the degree supported by the evidence; explain what could reverse the recommendation.
  • For uncertainty, define the cheapest next measurement or experiment with audience, signal, boundary and decision criterion without executing it.
  • Return results linked to the original requirement/hypothesis and observed deployment state.
  • Report data and causal limitations explicitly; do not transform lack of proof into proof of no effect.

Verdict

  • SUPPORTED: evidence supports the intended outcome within the stated population, window and causal limits.
  • NOT_SUPPORTED: valid evidence contradicts the declared outcome or violates a required guardrail.
  • INCONCLUSIVE: evidence cannot establish the outcome or causal interpretation.
  • BLOCKED: essential hypothesis, exposure identity or authorized data is unavailable.

Self-Check

  • Reconcile before returning. Check item-level evidence, requirement coverage, contradictions, scope, verdict, and applicable cleanup. Correct the report or authorized artifacts. Reuse valid evidence; do not automatically rescan the repository or rerun successful commands. Repeat checks only for relevant changes, failures, or unresolved evidence. Disclose remaining gaps.

Output Contract

Report in the user's language, in this order; label all five fields and state each fact once. Use controlled plain language: one fact per sentence, usually under 20 words, active voice, and one term per concept, with no synonyms for verdicts, IDs, or states. Small results may use one line per field; omit empty tables and do not copy linked artifacts:

  1. Result: The exact skill-specific verdict token first, then the supported outcome.
  2. Scope: Reviewed/changed scope, exclusions, baseline, and material assumptions.
  3. Evidence: Skill-specific fields below; distinguish facts, inferences, and unverified claims. Link artifacts; use tables when useful.
  4. Verification: Checks/results, unavailable evidence, and applicable cleanup/external state.
  5. Completion: Checklist: X/Y complete; Incomplete: None or each UNPROVEN item's reason, outcome impact, and exact next action; residual risks and required decisions.

Skill-specific evidence: Hypothesis, deployed exposure, baseline/control, metric definitions and quality, reproducible results and uncertainty, guardrails, causal limits, recommendation and next evidence action.

© levnikolaevich, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/operations-suite/skills/ln-72-product-outcome-evaluator of levnikolaevich/claude-code-skills.

Open the folder on GitHubat commit 0ce8796

Compare with similar skills

Ln 72 Product Outcome Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ln 72 Product Outcome Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ln 72 Product Outcome Evaluator this skilllevnikolaevich/claude-code-skills574—~1.9kAutomated safety check: PassMIT
Exploring LLM EvaluationsPostHog/posthog40k—~5.7kAutomated safety check: PassCustom licence
ObservabilityBuilderIO/agent-native7.1k—~7.3kAutomated safety check: PassNone
Langsmith ObservabilityOrchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT
RAG Observability Evalssickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: PassMIT
EvaluatorsArize-ai/phoenix12k—~1.7kAutomated safety check: PassCustom licence

Similar skills

  • Official

    Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment).

    40k GitHub stars~5.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Observability

    BuilderIO/agent-native

    Agent observability, evals, feedback, and experiments. An agent skill from BuilderIO/agent-native.

    7.1k GitHub stars~7.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Langsmith Observability

    Orchestra-Research/AI-Research-SKILLs

    LLM observability platform for tracing, evaluation, and monitoring.

    13k GitHub starsUsed in 2 repos~2.4k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Observability Evals

    sickn33/agentic-awesome-skills

    Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing.

    47k GitHub starsUsed in 2 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Evaluators

    Arize-ai/phoenix

    Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

    12k GitHub stars~1.7k tokensUpdated yesterday
    EducationAuto-check passed
  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 1 repo~8.1k tokens
    AI & LLM EngineeringAuto-check: notes

More from levnikolaevich/claude-code-skills

All 31 skills in this repo
  • Ln 53 Documentation Auditor

    levnikolaevich/claude-code-skills

    Audits documentation and comments for trust, coverage, consistency and freshness; read-only.

    574 GitHub stars~3.8k tokensUpdated 6 days ago
    Auto-check passed
  • Ln 81 Skill Reviewer

    levnikolaevich/claude-code-skills

    Reviews skill instructions, trigger boundaries and distribution contracts; not product code.

    574 GitHub stars~3.5k tokensUpdated 6 days ago
    Auto-check passed
  • Ln 11 Opportunity Evaluator

    levnikolaevich/claude-code-skills

    Evaluates new product opportunities through demand, channels and economics before committing to build.

    574 GitHub stars~3k tokensUpdated 6 days ago
    Auto-check passed
  • Ln 12 Product Requirements Builder

    levnikolaevich/claude-code-skills

    Defines product requirements, business rules and acceptance criteria for a committed intent; edits product docs only.

    574 GitHub stars~1.9k tokensUpdated 6 days ago
    Auto-check passed
  • Ln 13 Interaction Design Builder

    levnikolaevich/claude-code-skills

    Designs user flows, interaction states and mockups for a defined product scope; does not implement UI code.

    574 GitHub stars~1.8k tokensUpdated 6 days ago
    Auto-check passed
  • Ln 21 System Design Baseline Builder

    levnikolaevich/claude-code-skills

    Defines measurable architecture drivers and constraints before system design; edits architecture docs only.

    574 GitHub stars~2.5k tokensUpdated 6 days ago
    Auto-check passed

Questions about Ln 72 Product Outcome Evaluator

What does Ln 72 Product Outcome Evaluator do?

Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment. Ln 72 Product Outcome Evaluator is an agent skill from levnikolaevich/claude-code-skills. Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment.

How do I install Ln 72 Product Outcome Evaluator in Claude Code?

Run `npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a claude-code`. Or copy the skill folder (plugins/operations-suite/skills/ln-72-product-outcome-evaluator in levnikolaevich/claude-code-skills) into .claude/skills/ln-72-product-outcome-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Ln 72 Product Outcome Evaluator in Codex?

Run `npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a codex`. Or copy the skill folder (plugins/operations-suite/skills/ln-72-product-outcome-evaluator in levnikolaevich/claude-code-skills) into .agents/skills/ln-72-product-outcome-evaluator in your project. Codex loads it when a task matches its description.

Can I use Ln 72 Product Outcome Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ln-72-product-outcome-evaluator, .gemini/skills/ln-72-product-outcome-evaluator, .github/skills/ln-72-product-outcome-evaluator and .opencode/skills/ln-72-product-outcome-evaluator in your project.

What does Ln 72 Product Outcome Evaluator need to run?

SKILL.md names no scripts, command-line tools or credentials: Ln 72 Product Outcome Evaluator is instructions for the agent only.

Does Ln 72 Product Outcome Evaluator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ln 72 Product Outcome Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ln 72 Product Outcome Evaluator use?

Ln 72 Product Outcome Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ln 72 Product Outcome Evaluator use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ln 72 Product Outcome Evaluator?

Skills that share tags, products or a category with Ln 72 Product Outcome Evaluator: Exploring LLM Evaluations (PostHog/posthog, 40k stars), Observability (BuilderIO/agent-native, 7.1k stars), Langsmith Observability (Orchestra-Research/AI-Research-SKILLs, 13k stars) and RAG Observability Evals (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ln 72 Product Outcome Evaluator?

levnikolaevich (a GitHub user) maintains it in levnikolaevich/claude-code-skills, which has 574 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 5, 2026.

Source: levnikolaevich/claude-code-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.