Exploring LLM Evaluations
PostHog/posthog
Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment).
Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment.
$ npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install levnikolaevich/claude-code-skills ln-72-product-outcome-evaluator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/levnikolaevich/claude-code-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/operations-suite/skills/ln-72-product-outcome-evaluator .claude/skills/ln-72-product-outcome-evaluator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ln-72-product-outcome-evaluator" agent skill from https://github.com/levnikolaevich/claude-code-skills/tree/master/plugins/operations-suite/skills/ln-72-product-outcome-evaluator into .claude/skills/ln-72-product-outcome-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ln-72-product-outcome-evaluator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/levnikolaevich/claude-code-skills/tree/master/plugins/operations-suite/skills/ln-72-product-outcome-evaluatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install levnikolaevich/claude-code-skills ln-72-product-outcome-evaluator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/levnikolaevich/claude-code-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/operations-suite/skills/ln-72-product-outcome-evaluator .agents/skills/ln-72-product-outcome-evaluator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ln-72-product-outcome-evaluator" agent skill from https://github.com/levnikolaevich/claude-code-skills/tree/master/plugins/operations-suite/skills/ln-72-product-outcome-evaluator into .agents/skills/ln-72-product-outcome-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ln-72-product-outcome-evaluator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install levnikolaevich/claude-code-skills ln-72-product-outcome-evaluator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/levnikolaevich/claude-code-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/operations-suite/skills/ln-72-product-outcome-evaluator .cursor/skills/ln-72-product-outcome-evaluator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ln-72-product-outcome-evaluator" agent skill from https://github.com/levnikolaevich/claude-code-skills/tree/master/plugins/operations-suite/skills/ln-72-product-outcome-evaluator into .cursor/skills/ln-72-product-outcome-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ln-72-product-outcome-evaluator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/levnikolaevich/claude-code-skills.git --path plugins/operations-suite/skills/ln-72-product-outcome-evaluator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install levnikolaevich/claude-code-skills ln-72-product-outcome-evaluator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/levnikolaevich/claude-code-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/operations-suite/skills/ln-72-product-outcome-evaluator .gemini/skills/ln-72-product-outcome-evaluator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ln-72-product-outcome-evaluator" agent skill from https://github.com/levnikolaevich/claude-code-skills/tree/master/plugins/operations-suite/skills/ln-72-product-outcome-evaluator into .gemini/skills/ln-72-product-outcome-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ln-72-product-outcome-evaluator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install levnikolaevich/claude-code-skills ln-72-product-outcome-evaluatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/levnikolaevich/claude-code-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/operations-suite/skills/ln-72-product-outcome-evaluator .github/skills/ln-72-product-outcome-evaluator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ln-72-product-outcome-evaluator" agent skill from https://github.com/levnikolaevich/claude-code-skills/tree/master/plugins/operations-suite/skills/ln-72-product-outcome-evaluator into .github/skills/ln-72-product-outcome-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ln-72-product-outcome-evaluator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install levnikolaevich/claude-code-skills ln-72-product-outcome-evaluator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/levnikolaevich/claude-code-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/operations-suite/skills/ln-72-product-outcome-evaluator .opencode/skills/ln-72-product-outcome-evaluator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ln-72-product-outcome-evaluator" agent skill from https://github.com/levnikolaevich/claude-code-skills/tree/master/plugins/operations-suite/skills/ln-72-product-outcome-evaluator into .opencode/skills/ln-72-product-outcome-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ln-72-product-outcome-evaluator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ln-72-product-outcome-evaluatorEvaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment.
Ln 72 Product Outcome Evaluator is an agent skill from levnikolaevich/claude-code-skills. Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: Help your AI agent finish the job: solve the right problem, keep changes focused, and show what was verified. For Claude Code and Codex. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0ce8796. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ln 72 Product Outcome Evaluator loads about 1.9k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 905 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from levnikolaevich/claude-code-skills at commit 0ce8796, republished under its MIT licence (© levnikolaevich). 905 words, ~1,854 tokens.
.claude/skills/ln-72-product-outcome-evaluator/SKILL.md (or your agent's skills folder).Goal: Determine what available evidence supports about a delivered product outcome and recommend continuation, adjustment or stopping. Remain read-only: do not change instrumentation, experiments, user treatment, campaigns or product files.
Execution contract: The checklist defines completion. Track each item internally as PENDING, PROVEN with evidence, CLEARED with evidence its condition is absent, or UNPROVEN with a gap; reading, delegation, tool failure, a zero exit status, or a self-reported success is not proof; only the observed outcome is. Reconcile after each section. Before returning, resolve all PENDING, count only PROVEN and CLEARED, and apply verdict and approval rules to every gap.
Preserve intent, scope, and existing authorization. Continue authorized work; ask only for consequential unresolved choices or required external approval. When no one can answer during the run, state the exact question and apply the skill's verdict for the remaining gap instead of waiting or guessing. Scale depth to material risk without skipping checks. Preserve dependency and safety order; otherwise choose an appropriate verification method.
Accept equivalent user or repository evidence; no other skill, named artifact, or complete lifecycle is required. Preserve source requirement and decision IDs. Bind reused evidence to relevant source versions, dirty changes, configuration, and environment; invalidate only affected claims.
On continuation, reconcile task, authorization, current state, and unresolved evidence. For long work, return a compact continuation record or update an already authorized artifact; read-only skills do not persist it. Distinguish artifact readiness, verified behavior, and external-action authority.
Prepare authorized work before required approval. If blocked by an instruction, cite its exact source and unresolved boundary; do not invent approval gates from caution.
| Need | Preferred capability | Fallback |
|---|---|---|
| Original hypothesis | Product intent, baseline, experiment/measurement plan and accepted targets | Reconstruct from attributable sources; keep missing targets unknown |
| Outcome evidence | Authorized analytics, experiment results, customer behavior and cost/support evidence | Sanitized exports with explicit measurement limits |
| Analysis | Reproducible queries/statistics appropriate to the study design | Transparent arithmetic and qualitative inference; no fabricated causal confidence |
SUPPORTED: evidence supports the intended outcome within the stated population, window and causal limits.NOT_SUPPORTED: valid evidence contradicts the declared outcome or violates a required guardrail.INCONCLUSIVE: evidence cannot establish the outcome or causal interpretation.BLOCKED: essential hypothesis, exposure identity or authorized data is unavailable.Report in the user's language, in this order; label all five fields and state each fact once. Use controlled plain language: one fact per sentence, usually under 20 words, active voice, and one term per concept, with no synonyms for verdicts, IDs, or states. Small results may use one line per field; omit empty tables and do not copy linked artifacts:
Checklist: X/Y complete; Incomplete: None or each UNPROVEN item's reason, outcome impact, and exact next action; residual risks and required decisions.Skill-specific evidence: Hypothesis, deployed exposure, baseline/control, metric definitions and quality, reproducible results and uncertainty, guardrails, causal limits, recommendation and next evidence action.
© levnikolaevich, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/operations-suite/skills/ln-72-product-outcome-evaluator of levnikolaevich/claude-code-skills.
Open the folder on GitHubat commit 0ce8796
Ln 72 Product Outcome Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ln 72 Product Outcome Evaluator this skilllevnikolaevich/claude-code-skills | 574 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Exploring LLM EvaluationsPostHog/posthog | 40k | — | ~5.7k | Automated safety check: Pass | Custom licence | |
| ObservabilityBuilderIO/agent-native | 7.1k | — | ~7.3k | Automated safety check: Pass | None | |
| Langsmith ObservabilityOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | |
| RAG Observability Evalssickn33/agentic-awesome-skills | 47k | 2 repos | ~3.1k | Automated safety check: Pass | MIT | |
| EvaluatorsArize-ai/phoenix | 12k | — | ~1.7k | Automated safety check: Pass | Custom licence |
PostHog/posthog
Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment).
BuilderIO/agent-native
Agent observability, evals, feedback, and experiments. An agent skill from BuilderIO/agent-native.
Orchestra-Research/AI-Research-SKILLs
LLM observability platform for tracing, evaluation, and monitoring.
sickn33/agentic-awesome-skills
Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing.
Arize-ai/phoenix
Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
levnikolaevich/claude-code-skills
Audits documentation and comments for trust, coverage, consistency and freshness; read-only.
levnikolaevich/claude-code-skills
Reviews skill instructions, trigger boundaries and distribution contracts; not product code.
levnikolaevich/claude-code-skills
Evaluates new product opportunities through demand, channels and economics before committing to build.
levnikolaevich/claude-code-skills
Defines product requirements, business rules and acceptance criteria for a committed intent; edits product docs only.
levnikolaevich/claude-code-skills
Designs user flows, interaction states and mockups for a defined product scope; does not implement UI code.
levnikolaevich/claude-code-skills
Defines measurable architecture drivers and constraints before system design; edits architecture docs only.
Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment. Ln 72 Product Outcome Evaluator is an agent skill from levnikolaevich/claude-code-skills. Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment.
Run `npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a claude-code`. Or copy the skill folder (plugins/operations-suite/skills/ln-72-product-outcome-evaluator in levnikolaevich/claude-code-skills) into .claude/skills/ln-72-product-outcome-evaluator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a codex`. Or copy the skill folder (plugins/operations-suite/skills/ln-72-product-outcome-evaluator in levnikolaevich/claude-code-skills) into .agents/skills/ln-72-product-outcome-evaluator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ln-72-product-outcome-evaluator, .gemini/skills/ln-72-product-outcome-evaluator, .github/skills/ln-72-product-outcome-evaluator and .opencode/skills/ln-72-product-outcome-evaluator in your project.
SKILL.md names no scripts, command-line tools or credentials: Ln 72 Product Outcome Evaluator is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ln 72 Product Outcome Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ln 72 Product Outcome Evaluator: Exploring LLM Evaluations (PostHog/posthog, 40k stars), Observability (BuilderIO/agent-native, 7.1k stars), Langsmith Observability (Orchestra-Research/AI-Research-SKILLs, 13k stars) and RAG Observability Evals (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
levnikolaevich (a GitHub user) maintains it in levnikolaevich/claude-code-skills, which has 574 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 5, 2026.
Source: levnikolaevich/claude-code-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.