---
name: hire-or-pass
description: "Turn a hiring debrief into a typed decision — hire, hire at a different level, get more signal, or pass — from rubric evidence only, with probabilities and confidence. Use when asked should we hire this candidate, run the debrief, hire or no-hire, or when interviewers disagree and the manager has to call it. Produces the evidence state per competency, the four-way decision with probabilities, the act-or-hold verdict, and when holding, the exact competency to test and who tests it. Decision support for a human decision — a person owns the hire."
version: 1.0.0
---

# Hire or Pass

Debriefs go wrong when the loudest interviewer wins or when everyone waits to see what the manager thinks. This skill forces the evidence into a fixed shape — per competency, strongest example and strongest counter-example, from whom — and defines the four possible outcomes before the discussion starts. The output is a distribution over those outcomes and, when it is close, the *specific* extra signal to go get.

**Boundary.** This is decision support. A person makes and owns the hiring decision, applies the company's process and the law where they are, and is accountable for it. Never let a distribution replace that.

## Where this sits

After [`hiring-rubric`](../hiring-rubric/SKILL.md) (which defines the competencies) and [`engineering-hiring-rubric`](../engineering-hiring-rubric/SKILL.md), and after the interviews. This one runs the debrief decision.

## What This Skill Produces

- **The evidence state** — per must-have competency: strongest example, strongest counter-example, which interviewer saw it, rubric scores
- **The decision** — `hire` / `hire_at_different_level` / `more_signal` / `pass` with probabilities and confidence
- **Act or hold** — thresholds applied; a close hire/pass is a `more_signal` by definition
- **The signal plan** — when holding: the competency to test, the format, who runs it, by when
- **The debrief record** — one paragraph for the file: evidence summary, decision, who decided

## Required Inputs

Ask for these; do not infer from "vibes" or an interviewer's overall thumbs:
- **role** — title, level, the 3–5 must-have competencies from the rubric
- **evidence** — per competency: strongest observed example and strongest counter-example, with the interviewer who saw it
- **scores** — rubric scores per interviewer per competency
- **references** — if taken: confirmed or contradicted what
- **risks** — gaps, level mismatch, red flags *with the evidence behind them*
- **process** — any missed interviews or untested competencies

## Framework: Four Defined Outcomes

| Option | It applies when |
|---|---|
| `hire` | Every must-have competency has positive evidence from at least two interviewers, and no red flag has evidence behind it |
| `hire_at_different_level` | The evidence supports the competencies at a different level than the role was opened at |
| `more_signal` | A must-have was untested, or the evidence is split between interviewers — one targeted extra conversation decides it |
| `pass` | A must-have competency has negative evidence, or a red flag is substantiated |

**Thresholds.** Act when confidence ≥ 0.65 and the winning option's probability ≥ 0.7. A close `hire` / `pass` **is** a `more_signal` outcome: name the competency and who tests it. Never decide on the distribution alone.

**Evidence, not impressions.** "Strong communicator" is an impression. "Explained the outage timeline to a non-engineer in the panel without being asked — seen by R." is evidence. Only evidence enters the state.

**Getting probabilities.** Any calibrated decision model that returns a probability per defined option serves this contract; the repo ships an adapter (`node integrations/jev/decide.mjs hire-or-pass --state state.json`, contract in `integrations/jev/decisions/`). Without one, the assistant estimates and labels it. Either way the decision is the hiring manager's.

## Output Format

### Debrief decision: [role, level] · candidate [initials]

**Evidence by competency**
| Competency | Strongest example (who) | Strongest counter (who) | Scores |
|---|---|---|---|
| … | … | … | … |

**Decision**
| Option | Probability |
|---|---|
| hire | 0.xx |
| hire_at_different_level | 0.xx |
| more_signal | 0.xx |
| pass | 0.xx |

Confidence: 0.xx · Source: [model / estimate]

**Verdict:** **[act: option]** or **[hold → more_signal: competency, format, owner, by when]**
**Debrief record:** [one paragraph — evidence summary, the decision, who made it]

## Quality Checks
- [ ] Every must-have has an evidence row; untested ones are marked and force `more_signal`
- [ ] Counter-examples are recorded, not just the good moments
- [ ] Red flags are listed with evidence or dropped
- [ ] All four probabilities shown; the verdict cites the thresholds
- [ ] The record names the human who decided
- [ ] Nothing in the state is a protected characteristic or a proxy for one

## Anti-Patterns
- **Overall thumbs** replacing per-competency evidence
- **Anchoring on the manager** — collect scores before anyone speaks
- **"Culture fit"** as an undefined competency
- **Deciding a coin flip** — 0.5 / 0.5 hire/pass is a `more_signal`, always
- **Letting the model hire** — it produces evidence about evidence; a person decides

## Example Trigger Phrases
- "Run the debrief — three interviewers, split panel."
- "Hire or pass on this candidate? Here are the scorecards."
- "We're at hire vs. pass 50/50 — what do we do?"
