Agent skill

Happier Instruction Eval

by happier-dev in happier-dev/happier

Evaluate two or more Happier instruction, constitution, or skill variants with blinded organic tasks, controlled context, behavior-based scoring, privacy-safe evidence, and an advisory synthesis.

MITAuto-check passed

Install Happier Instruction Eval

skills CLI
$ npx skills add happier-dev/happier --skill happier-instruction-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install happier-dev/happier happier-instruction-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/happier-dev/happier.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/happier-instruction-eval .claude/skills/happier-instruction-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
happier-instruction-eval
GitHub stars
1.9k
Token cost
~1.3k tokens
SKILL.md length
631 words
Files
1
Skills in repo
28
Repo updated
First seen
Licence
MIT

At a glance

Evaluate two or more Happier instruction, constitution, or skill variants with blinded organic tasks, controlled context, behavior-based scoring, privacy-safe evidence, and an advisory synthesis.

  • Works in 6 steps: Establish the decision → Control the comparison → Keep runners blind → …
  • Explicitly asks to compare/evaluate instruction variants
  • SKILL.md covers 1. Establish the decision, 2. Control the comparison, 3. Keep runners blind and 4. Measure behavior, not…, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Happier Instruction Eval is an agent skill from happier-dev/happier. Evaluate two or more Happier instruction, constitution, or skill variants with blinded organic tasks, controlled context, behavior-based scoring, privacy-safe evidence, and an advisory synthesis. Use only when the user explicitly asks to compare/evaluate instruction variants or an approved instruction program names an evaluation boundary.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Web, Desktop & Mobile client and orchestrator for Codex, Claude Code, OpenCode, Pi, Cursor, Grok, Antigravity, Kimi, Augment Code, Qwen, fully end-to-end encrypted. The licence is MIT.

When your agent uses it

  • Explicitly asks to compare/evaluate instruction variants
  • An approved instruction program names an evaluation boundary

Example prompts

  • “/happier-instruction-eval”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Establish the decision
  2. Control the comparison
  3. Keep runners blind
  4. Measure behavior, not self-report
  5. Judge under neutral labels
  6. Synthesize without automatic mutation

What it can do on your machine

Read from SKILL.md and the folder at commit 493820f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Happier Instruction Eval loads about 1.3k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 631 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from happier-dev/happier at commit 493820f, republished under its MIT licence (© happier-dev). 631 words, ~1,266 tokens.

Download SKILL.mdSave it as .claude/skills/happier-instruction-eval/SKILL.md (or your agent's skills folder).
name
happier-instruction-eval
description
Evaluate two or more Happier instruction, constitution, or skill variants with blinded organic tasks, controlled context, behavior-based scoring, privacy-safe evidence, and an advisory synthesis. Use only when the user explicitly asks to compare/evaluate instruction variants or an approved instruction program names an evaluation boundary.

Happier Instruction Evaluation

Evaluate whether instruction wording changes agent behavior without telling the agents they are being evaluated. This workflow is advisory: it produces evidence for a human decision and never edits the canonical instruction owner by itself.

1. Establish the decision

Name:

  • the instruction owner and variants being compared, including the current baseline;
  • the concrete behavior the change is meant to improve or failure it is meant to prevent;
  • one organic task or a small risk-selected set of tasks that can expose that difference;
  • a rubric of observable outcomes fixed before any run;
  • the user decision the evidence will inform.

Do not evaluate prose elegance in isolation. A useful task forces the instruction to affect routing, investigation, ownership, implementation shape, validation, stopping, or reporting. Skip evaluation when a source inspection or deterministic check can decide the question directly.

2. Control the comparison

Hold constant everything except the instruction variant when practical:

  • use the same task prompt, repository basis, allowed tools, permissions, time/effort budget, and available evidence;
  • give each run only the ordinary context an agent would receive for that task;
  • label variants and output locations neutrally so neither runner nor judge sees “baseline,” “preferred,” model identity, or another run's existence;
  • isolate writes in separate temporary directories or authorized worktrees; never switch, clean, reset, stash, or overwrite the primary shared checkout;
  • prevent external mutations, destructive actions, secrets, and sensitive-data access unless the user separately authorizes that exact evaluation surface.

Do not freeze or package a release representation. Temporary instruction variants and isolated outputs are test inputs, not release artifacts.

3. Keep runners blind

Each runner receives an organic-looking engineering request, not an evaluation brief. Do not mention the rubric, competing variants, expected lesson, or favored outcome. Do not ask the runner whether it followed the instruction.

Use minimal inherited conversation context. Never expose private transcripts, credentials, customer data, or unrelated work. Prefer synthetic tasks, public evidence, or bounded repository tasks. Historical conversations require explicit sensitive-data authorization and sanitization.

If an assigned tool or model differs between runs, record that as a confound rather than treating model agreement or disagreement as proof. Use repeated trials only when outcome variability is decision-material; never manufacture a fixed sample count.

Show full SKILL.md (265 more words)Show less

4. Measure behavior, not self-report

Inspect what each run actually did:

  • sources and instruction owners read;
  • questions asked versus empirical facts investigated;
  • canonical owner and affected corridor identified;
  • split-brains, unsupported requirements, or scope drift introduced or prevented;
  • edits and artifacts produced;
  • tests, live checks, and falsifiers actually run;
  • unsafe, irrelevant, or ceremonial work avoided;
  • final claims, uncertainty labels, and residual risk.

Score each rubric item from the artifacts and tool evidence. A polished explanation or claimed compliance is not evidence. Mark unavailable observations and confounds explicitly.

5. Judge under neutral labels

Give one judge all outputs under neutral labels and the same precommitted rubric. The judge must:

  1. score each outcome criterion independently;
  2. cite the behavior or artifact supporting each score;
  3. identify regressions, omissions, confounds, and ties;
  4. recommend retain, revise, combine, reject, or run one discriminating follow-up;
  5. avoid inferring variant identity, author intent, or model quality.

The orchestrator re-derives decision-material claims from the underlying artifacts. Agreement between runners, judge, and orchestrator raises a question's priority; it does not replace evidence.

6. Synthesize without automatic mutation

Report:

  • decision and tested behavior;
  • task and controlled basis;
  • anonymized rubric results with evidence pointers;
  • confounds and unobserved surfaces;
  • which wording or structural change earned its place and why;
  • the smallest recommended canonical-owner edit, or that no change is justified.

Use the final response unless the user requested durable evaluation tracking. Do not create a new report file, update AGENTS.md, or propagate a variant to another repository without explicit change authority. If an accepted change affects the 0.2 source line, use .agents/skills/happier-port-0-2-to-0-3 for its destination disposition.

© happier-dev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/happier-instruction-eval of happier-dev/happier.

Open the folder on GitHubat commit 493820f

Compare with similar skills

Happier Instruction Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Happier Instruction Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Happier Instruction Eval this skillhappier-dev/happier1.9k—~1.3kAutomated safety check: PassMIT
Eval-Driven Development Harnessaffaan-m/ECC274k—~1.5kAutomated safety check: PassMIT
Evalalirezarezvani/claude-skills28k1 repos~618Automated safety check: PassMIT
Eval Harnessaffaan-m/ECC274k—~2.2kAutomated safety check: PassMIT
Harness Evaltech-leads-club/agent-skills7k—~3.9kAutomated safety check: PassCC-BY-4.0
Eval Harnessaffaan-m/ECC274k1 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

    274k GitHub stars~1.5k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Eval

    alirezarezvani/claude-skills

    Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

    28k GitHub starsUsed in 1 repo~618 tokens
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

    274k GitHub stars~2.2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Harness Eval

    tech-leads-club/agent-skills

    Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports.

    7k GitHub stars~3.9k tokensUpdated 17 days ago
    Agent WorkflowsAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi

    274k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • OmniRoute CLI Evals

    diegosouzapw/OmniRoute

    Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

    74k GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from happier-dev/happier

All 28 skills in this repo
  • Happier Review

    happier-dev/happier

    Conduct evidence-backed Happier code, plan-completeness, session, worktree, feature, commit, branch, PR, codebase, and release-readiness reviews with affected-corridor analysis, high-confidence…

    1.9k GitHub stars~4.5k tokensUpdated today
    Auto-check passed
  • Happier CI Stabilize

    happier-dev/happier

    Stabilize failing, flaky, slow, or repeatedly rerun Happier CI and nightlies by collecting all reachable failures from one exact attempt, correcting canonical causes in one batch, simplifying…

    1.9k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Happier Commit Worktree

    happier-dev/happier

    Reconnoiter, classify, validate, group, and commit a large or continuously changing Happier worktree as coherent, human-understandable commits while preserving concurrent work and excluding…

    1.9k GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Happier Release

    happier-dev/happier

    Resolve Happier's private release authority and run an exact-SHA release or nightly through cheap admission, verified CI evidence, resumable immutable candidates, and terminal publication proof.

    1.9k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Happier Diagnose

    happier-dev/happier

    Diagnose and explain a Happier runtime, session, daemon, provider (Claude/Codex/OpenCode), authentication, or connectivity incident from logs, structured diagnostics, runtime state, and source…

    1.9k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Happier Implement

    happier-dev/happier

    Implement, change, build, fix, refactor, migrate, or apply accepted review findings in the Happier repositories with canonical-owner discovery, scope-preserving solution economy, TDD, efficient…

    1.9k GitHub stars~4.2k tokensUpdated today
    Auto-check passed

Questions about Happier Instruction Eval

What does Happier Instruction Eval do?

Evaluate two or more Happier instruction, constitution, or skill variants with blinded organic tasks, controlled context, behavior-based scoring, privacy-safe evidence, and an advisory synthesis. Happier Instruction Eval is an agent skill from happier-dev/happier. Evaluate two or more Happier instruction, constitution, or skill variants with blinded organic tasks, controlled context, behavior-based scoring, privacy-safe evidence, and an advisory synthesis.

When should I use Happier Instruction Eval?

Happier Instruction Eval fits situations like: explicitly asks to compare/evaluate instruction variants; an approved instruction program names an evaluation boundary.

How do I install Happier Instruction Eval in Claude Code?

Run `npx skills add happier-dev/happier --skill happier-instruction-eval -a claude-code`. Or copy the skill folder (.agents/skills/happier-instruction-eval in happier-dev/happier) into .claude/skills/happier-instruction-eval in your project. Claude Code loads it when a task matches its description.

How do I install Happier Instruction Eval in Codex?

Run `npx skills add happier-dev/happier --skill happier-instruction-eval -a codex`. Or copy the skill folder (.agents/skills/happier-instruction-eval in happier-dev/happier) into .agents/skills/happier-instruction-eval in your project. Codex loads it when a task matches its description.

Can I use Happier Instruction Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add happier-dev/happier --skill happier-instruction-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/happier-instruction-eval, .gemini/skills/happier-instruction-eval, .github/skills/happier-instruction-eval and .opencode/skills/happier-instruction-eval in your project.

What does Happier Instruction Eval need to run?

SKILL.md names no scripts, command-line tools or credentials: Happier Instruction Eval is instructions for the agent only.

Does Happier Instruction Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Happier Instruction Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Happier Instruction Eval use?

Happier Instruction Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Happier Instruction Eval use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Happier Instruction Eval?

Skills that share tags, products or a category with Happier Instruction Eval: Eval-Driven Development Harness (affaan-m/ECC, 274k stars), Eval (alirezarezvani/claude-skills, 28k stars), Eval Harness (affaan-m/ECC, 274k stars) and Harness Eval (tech-leads-club/agent-skills, 7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Happier Instruction Eval?

happier-dev (a GitHub organization) maintains it in happier-dev/happier, which has 1,876 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 7, 2026.

Source: happier-dev/happier on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.