Diagnosing Superpowers Sessions
obra/superpowers
Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.
Diagnose why a scored answer failed and who owns the fix, then cluster the failures by shared root cause.
$ npx skills add malloydata/publisher --skill eval-diagnose -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install malloydata/publisher eval-diagnose --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-diagnose .claude/skills/eval-diagnose && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-diagnose" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-diagnose into .claude/skills/eval-diagnose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-diagnose", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/malloydata/publisher/tree/main/skills/eval-diagnoseType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add malloydata/publisher --skill eval-diagnose -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install malloydata/publisher eval-diagnose --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-diagnose .agents/skills/eval-diagnose && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-diagnose" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-diagnose into .agents/skills/eval-diagnose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-diagnose", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-diagnose -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install malloydata/publisher eval-diagnose --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-diagnose .cursor/skills/eval-diagnose && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-diagnose" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-diagnose into .cursor/skills/eval-diagnose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-diagnose", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/malloydata/publisher.git --path skills/eval-diagnose--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add malloydata/publisher --skill eval-diagnose -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install malloydata/publisher eval-diagnose --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-diagnose .gemini/skills/eval-diagnose && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-diagnose" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-diagnose into .gemini/skills/eval-diagnose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-diagnose", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install malloydata/publisher eval-diagnoseInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add malloydata/publisher --skill eval-diagnose -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-diagnose .github/skills/eval-diagnose && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-diagnose" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-diagnose into .github/skills/eval-diagnose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-diagnose", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-diagnose -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install malloydata/publisher eval-diagnose --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-diagnose .opencode/skills/eval-diagnose && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-diagnose" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-diagnose into .opencode/skills/eval-diagnose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-diagnose", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-diagnoseDiagnose why a scored answer failed and who owns the fix, then cluster the failures by shared root cause.
Eval Diagnose is an agent skill from malloydata/publisher. Diagnose why a scored answer failed and who owns the fix, then cluster the failures by shared root cause. Walk dataset, agent-call, getcontext/model, getcontext/retrieval, construction, then model-definition. Append issue events to the file ledger, linked by traceId, one per cluster. Use after eval-answer, when triaging a run, or before changing a model. Does not edit the model (eval-improve).
Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts (for example `reference/output-contract.md`, `scripts/cluster_failures.py` and `scripts/cluster_failures_test.py`).
The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 39a546f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 4 files in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Diagnose loads about 5.9k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 3,554 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from malloydata/publisher at commit 39a546f, republished under its MIT licence (© malloydata). 3,554 words, ~5,874 tokens.
.claude/skills/eval-diagnose/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Consumes a score event from skill:eval-answer and answers: why did this fail,
and who owns the fix?
Scope boundary: write the diagnosis before any edit exists. This skill never edits a model and never proposes a patch beyond naming the gap. Diagnosis that is allowed to edit becomes justification for an edit somebody already wanted.
Do not diagnose a contaminated attempt or an environment failure. Those are harness or ops, not model work.
Which cases. no_match on the dev split, which is what diagnose.py
selects by default. One exception, and it needs two arms to earn it: a case that
came back near_match in BOTH runs of a pair is not the judge hedging, it is
the model failing to distinguish two readings the question does, and no rubric
repair closes that. Take the stable list flip_table.py prints and pass
--only <qids> --verdicts near_match. Never diagnose a one-armed near_match;
that is noise, and it sends an agent to fix a model that is already right.
A correct answer can still carry a finding. A case that answered right
while a required entity never reached it is diagnosed too, for the retrieval
miss alone, and diagnose.py selects it automatically. The answer was right by
another route, and naming that route is the job: on the run this rule comes
from it was always the same one, the agent rebuilding the model's own measure
inline. That held while the measure was count() and failed the moment one
carried a grain rule, producing the run's only wrong answer. Three of the four
findings in that run sat on passing cases and, before this, produced nothing.
"It worked anyway" is not a reason to leave the model or the skills
unfixed. An answer that is right without the model's own entity is right for
now, not right by design. Write the issue against the miss, record that the
answer was correct so nobody reads it as a wrong number, and do not soften the
finding because the number came out right. --no-retrieval-misses opts out for
a run that only wants answer failures.
Holdout is withheld so the acceptance check keeps something the improve step
never saw. A measure-only run never reaches improve, so it is holding those
cases back from nothing: pass --include-holdout there. The script refuses it
on a run that already carries a candidate, because that run's holdout is the
only thing left that can falsify the edit.
Say what the clusters do not cover. diagnose.py prints a coverage account
of every non-passing case and which bucket it fell in -- holdout, contaminated,
a verdict outside --verdicts, unscored, or selected and not diagnosed. Quote
it whenever you report clusters. One run's six clusters were read as covering
its failures; they covered 8 of 18, and each individual exclusion had been
correct and added up nowhere.
Walk in this order and stop at the first with positive evidence. A later
label requires ruling out the earlier ones. Write component with these strings,
never "C1" / "C2" / "C3":
component | Question |
|---|---|
dataset | Bad question, bad or missing golden, or environment drift? |
agent-call | Did the agent ask for the needed concepts, with the right type and scope? |
get_context/model | Is the needed entity absent, undocumented, weakly labeled, duplicated, or missing guidance? |
get_context/retrieval | Was an on-target, in-scope request against a well-described entity ranked or grouped wrong? Check the call's scopes first: a call pinned to one source cannot return another source's entity, and that miss is agent-call. |
construction | Did sufficient context arrive, and the agent still built the wrong query? |
model-definition | Is a measure, join, filter convention, or source semantically wrong? |
owner is separate: model, retrieval, agent-skill, or dataset. There
is no environment owner: an environment failure stops the run before
diagnosis (see the boundary above), so no issue can carry it.
If the model's documentation were perfect, would the agent do the right thing?
agent-skill. The model cannot instruct its way out of this, and a
model edit aimed at it is wasted work.model.dataset.Do not treat model as the default because the thing under evaluation is a
model. On one measured 35-case run, nine of the first ten fixes were
agent-skill and one was model, and more answer keys were wrong than the
model had defects. Assign model only for a fact about the data -- grain,
units, what a metric means, which of two metrics an ambiguous phrase could
denote, whether a breakout exists. Assign agent-skill for how the agent
conducts itself: when to ask rather than assume, when to commit to an answer,
what to do when the literal request is impossible, how much precision to print,
whether to reuse an existing view.
A model doc cannot override a skill instruction. This is the trap the
question above exists to catch. Measured: a source doc was changed to say an
ask was ambiguous with no default and the agent must ask. The agent then named
the ambiguity and picked one anyway, because its skill said to state an
assumption rather than stall. Six cases turned on it, and none moved until the
skill was changed. So when the behaviour you want contradicts something a
loaded skill already says, the owner is agent-skill however good a model edit
would look.
Two shapes accounted for every agent-skill defect in that run, and both are
worth testing a candidate rule against:
Count Skill invocations in the answerer transcripts for the run. Measured on
one 35-case run with a 12-skill manifest: malloy-analysis loaded 34 times,
malloy-charts once, the other ten zero times -- including two that
malloy-analysis tells the agent outright to load, by reference, before it
writes a query. Cross-skill references do not reliably fire. A recommendation
to edit a file the agent never opens is not actionable, so name the file the
transcripts show it reading.
construction requires proving the needed entities and governing guidance were
in the returned context. A server trace proves what Publisher returned, not what
the host kept after compaction. If the rendered tool response is gone, mark
sufficiency unknown and do not assign construction.
Always report construction eligibility as eligible / total. That is a
diagnostic conditional, not a causal comparison.
For each get_context call, load the stored retrieval trace by the traceId on the
tool_call event. Write down, before you interpret anything:
rankedSummary on the attempt's tool_call events (its targets list
carries per-target ranks); the full trace body is behind your host's
trace lookup.
Count from the trace, never from recollection.Resolve aliases to the real source (join_one: bldg is fac_building uses
fac_building). Count ranks from the trace, not from recollection.
The needed set comes from golden metadata or from the entities the corrected answer required. Do not invent it from the question's nouns alone.
Presence leads; rank refines. Every needed entity present with a wrong
answer is prima facie construction. A needed entity that never appeared is
never construction, no matter how wrong the query looks. Everything
present but buried deep under noise the agent reasonably skipped is
get_context/retrieval once the request itself was on-target.
Use these codes verbatim. Re-wording them destroys the cross-answer pattern.
| Code | When | Owner |
|---|---|---|
BAD-REFERENCE | the golden itself is wrong, and you can name the defect | dataset |
AMBIGUOUS-REFERENCE | the key is untrustworthy as a score, but a replacement is not uniquely determined (two honest replays disagree; later cases may confirm a convention) | dataset |
BAD-QUESTION | the question is unanswerable as written, or underspecified (ties, rank without order) | dataset |
CORRECT-SUPERSET | every expected row present, plus extra context | none (it passed) |
A cheap tell for a bad golden: impossible magnitude; identical values across
entities that should differ; SUM / COUNT(*) over a join that duplicates on
both sides. AVG / STDDEV / MIN / MAX survive uniform duplication, so
fanout alone proves nothing.
BAD-REFERENCE and AMBIGUOUS-REFERENCE are first-class outcomes, not awkward
misses. Goldens often encode assumptions we want in the model. They can also
be wrong. Flag the case when the key is defective or when two justified
replays disagree: do not edit the model to match a bad or unsettled key, and
do not invent a replacement number. BAD-REFERENCE goes to Repair a bad
golden. AMBIGUOUS-REFERENCE goes to Hold an ambiguous golden. Do not leave
the run looking like the model failed.
This skill classifies and hands off. Write the issue with
owner: dataset and stop for that case. The conductor (skill:eval-loop)
either repairs the golden or holds it as ambiguous. Do not capture a
replacement golden from inside diagnosis if you are not also conducting; a
diagnosis that writes a new key without a version bump silently changes what
earlier scores meant.
Prior score events are not rewritten. They keep the old golden_revision.
| Code | When | Owner |
|---|---|---|
NEVER-ASKED | no utterance targeted a needed concept | agent-skill, and model if nothing would have prompted the ask |
VAGUE | compound or generic utterances, so nothing could rank | agent-skill |
QUESTION-VOCAB | utterances parroted the question where the data uses other words | agent-skill, and model if that vocabulary is undocumented |
RULE_UNWRITTEN | the data is present and the model does not encode the rule for combining or filtering it, so whoever answers invents one. Qualified (guessable) when the data forces the rule -- list price minus sale price -- and (arbitrary) when it is a business decision nobody could derive, such as a season window or whether revenue is net of tax. Only the stated conventions tell those apart: from the model alone the two look identical, and an unqualified reading defaults to the harmless one | model: write the rule down |
ASSUMED | assumed a scope or convention instead of checking | model if nothing warned; agent-skill otherwise |
WRONG-TYPE-OR-SCOPE | asked, but with the wrong target type or an empty/wrong scope | agent-skill |
If the agent could not reasonably have known to ask, that is a model gap.
| Code | When | Owner |
|---|---|---|
MISSING | no query over this model could produce the concept. Not merely "no named measure": if the parts are present and only the formula is absent, that is RULE_UNWRITTEN | model: add an entity, a dimension value, or a data source |
NOT-RETURNED | it exists, the ask was on target, it never came back | model: labels, docs, synonyms, index |
LOW-RANK | returned, buried under noise the agent reasonably skipped | model |
AMBIGUOUS | several candidates and nothing says which this question means. Judged over the entities AND their docs: a #(doc) that resolves the choice makes it MODELLED, which is why a sentence of documentation is usually the whole fix | model: "use X for …, Y when …" |
GUIDANCE-NOT-RETRIEVED | entities came back, governing guidance did not | model: put guidance on the entities agents search for |
GUIDANCE-DECLINED | guidance was retrieved and judged inapplicable | model: state the business default, not a caveat |
Before diagnosing a NOT-RETURNED, check the call actually returned
nothing. The tool_call event carries retrieval_mode and rankedSummary.
A rankedSummary of null means the response could not be read, not that it
was empty -- a body too large for the model's context is written to a file, and
the run records no summary for it. There is nothing to diagnose there: say the
call was not measured.
And never attribute a miss to the embedding index unless retrieval_mode
says indexing or error. The mode is recorded per call precisely so this is checkable.
A diagnosis once explained a phantom empty result with "the index was not ready
yet" on a run whose index was ready before the first question and whose
response did hold results; the empty list was a parsing bug in the harness. An
unverifiable cause that sounds right is worse than needs_human, because it
closes the finding.
A missing join is coverage, not an agent-call miss. The model has to volunteer
relationships. A declared join is not a retrieval entity; do not look for it in
get_context results.
| Code | When | Owner |
|---|---|---|
RETRIEVAL | model looks right, utterance on target, rank or grouping still failed | retrieval |
Prove it before you use this code: search a distinctive phrase from the entity's
own doc. If a rare token retrieves it and ordinary phrasing does not, say so
with both queries. Otherwise it is still NOT-RETURNED / LOW-RANK.
| Code | When | Owner |
|---|---|---|
WRONG-PICK | needed entity returned, used a different one | model if indistinguishable; agent-skill if docs distinguished them |
SCOPE | right entities, wrong population | model if the scope rule was undocumented |
GRAIN | right entities, wrong grain | model or agent-skill |
FILTER-LITERAL | filter literal did not match stored values | model (document the stored form) and agent-skill |
UNDERSPECIFIED | the QUESTION is unclear, not the model -- "adjusted sales" that never says what the adjustment is, or a business that has not decided whether revenue includes tax | the asker, or whoever owns the definition. Never scored against the model |
SYNTAX | could not express it; execute errors; never submitted | agent-skill |
Use when the entity was found and used, and the definition or the data behind it is wrong (bad grain, wrong join key, inverted filter). Owner: model. A doc whose factual claim the data contradicts (a population statement, a grain claim) is also model-definition: the SQL may be right while the stated contract is false, and an agent that trusts the doc answers wrongly without ever failing a query. Probe the claim before writing the issue.
| Signature | Look here |
|---|---|
| Extremes match, means do not | Population, filter, or join scope |
| Same row count, values differ | Wrong column or literal, not joins |
| Row count differs, all expected rows present | Superset; often not an error |
| Right keys, wrong aggregates on a minority | Undeclared or wrong-cardinality relationship |
| Off by a clean integer multiple | Fanout; which side of the join is non-unique |
| Zero errors, few calls, fast, confidently wrong | The model steered it |
| Identical high-precision values across entities that should differ | Cross-contamination join |
| A magnitude that cannot be true | Fanout, possibly in the golden |
Mine the agent's prose, not only its calls. It often names the gap.
Every signature above is about reading the ANSWER. They apply just as much to your own probes, which is the next section, and the fanout rows apply hardest: a probe is a query you wrote in a hurry against a model you have just met.
Evidence you generate yourself is not privileged over evidence you are handed. A real finding reported that a model's own documented recipe produced an impossible cumulative percentage, over 100% partway through the series, and cited a direct probe as proof. Re-running the recipe against the source the documentation actually routes to gave a textbook result: correct row count, monotonic, exactly 100% at the final point. The probe had been run against the PARENT source, and a measure summed across the dimensions the derived source exists to pin fans out. The tell was already in the diagnoser's own numbers: absolute counts orders of magnitude beyond any possible population. The ratio still looked well behaved, because fanout cancels top and bottom.
Three requirements, before a probe becomes a finding:
A diagnose pass at this precision is a lead generator, not a verdict. Every model-owned finding deserves a probe of its own before it justifies an edit.
Append to <workdir>/runs/<runId>/events.jsonl with kind: issue
(shapes in skill:eval-answer reference/ledger-schema.md):
issue_id, affected qids, primary_code, contributing_codescomponent, owner, severity, confidencesufficiency (sufficient / insufficient / unknown)traceIds, not copied trace payloadsdiagnosis: the suspected shared entity, file, or root cause, written
before any edit existsThen issue_status with status: open. Status is always an event. Readers
take the latest issue_status for that issue_id.
The issue backlog in the event log is the output, not per-question prose. Diagnose reads dev cases only; a holdout case with a bad score stays undiagnosed so the acceptance check keeps something the improve step never saw.
Eight failures are rarely eight problems. They are more often two or three,
each surfacing in several cases, and the edit worth making is the one that
clears a group. So the unit handed to skill:eval-improve is the cluster, not
the case, and one issue event covers all of its cases rather than one per case.
Group on shared cause, not shared symptom. Two cases that both returned a
wrong revenue number belong together only if the same entity, doc gap, or
convention explains both. Same owner and same component is a hint, never a
criterion: two MISSING issues about different missing entities are two
clusters, and merging them produces an edit that fixes neither cleanly.
Order clusters by how many cases they would fix. Cluster the non-model owners
too, in their own clusters, so nothing is lost on the way to the backlog --
but keep them separate, because only owner: model may proceed to an edit.
Say what you considered merging and chose not to. A cluster is a claim that one change fixes N cases, and the near-misses are what a reviewer needs to falsify it.
Falsify a behavioural cluster against the passes. This skill reads failures
only, so any behaviour common to the whole run looks causal from inside it.
diagnose.py hands the clustering step a CONTROLS block: the same
measurements -- retrieval calls, targets carrying no search_text, queries,
skills opened, turns -- taken on the cases that PASSED. Before claiming a
behaviour explains a cluster, compare it there. If it occurs at a similar rate
in the passes it does not separate the groups: mark the cluster contributing
rather than primary and do not route it to an edit as the root cause.
Measured, on the run this comes from: the largest cluster said the agent
"substitutes broad enumeration for targeted retrieval", and bare targets were
25% of all targets in the failures against 26% in the passes. What actually
separated them was volume -- failures made about 50% more retrieval calls --
and question difficulty explains that at least as well as call style does. With
no controls at all, a behavioural cluster is unfalsified, which is a different
claim from confirmed.
Still no patch. Naming the shared root cause precisely enough that someone else
can design the edit is the whole job here; the edit itself is
skill:eval-improve, working from this. A cluster carrying an honest open
question is more useful than one carrying a remedy nobody probed.
Only owner: model proceeds to eval-improve. Skill findings go back into
the analysis or phrase-detection skill. Retrieval findings go to the tool.
BAD-REFERENCE and AMBIGUOUS-REFERENCE go to the golden side door in
skill:eval-loop (repair or hold). Do not send them to improve. Other dataset
findings (a bad question, a case worth excluding) go back to the case in
cases.jsonl via the conductor. Routing a skill bug into the model is
how models accumulate scar tissue.
construction when sufficiency is unknown.skill:eval-answer: the score this consumes.skill:eval-improve: smallest model edit, owner: model only.skill:eval-loop: golden hold/repair, the acceptance check, and checkpoint.© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts) in skills/eval-diagnose of malloydata/publisher.
Open the folder on GitHubat commit 39a546f
Eval Diagnose next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Diagnose this skillmalloydata/publisher | 116 | — | ~5.9k | Automated safety check: Pass | MIT | |
| Diagnosing Superpowers Sessionsobra/superpowers | 296k | 3 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Code Design Rationale Investigatorcursor/plugins | 10k | 9 repos | ~2.6k | Automated safety check: Pass | None | |
| Pester Failure AnalysisPowerShell/PowerShell | 56k | — | ~5.1k | Automated safety check: Pass | MIT | |
| OpenLogi macOS Permissions TriageAprilNEA/OpenLogi | 23k | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Paseo Committeegetpaseo/paseo | 20k | 1 repos | ~496 | Automated safety check: Pass | Custom licence |
obra/superpowers
Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.
cursor/plugins
Digs into why code is shaped the way it is by checking git history, pull requests and connected tools in parallel, then reporting a cited read on the tradeoffs.
PowerShell/PowerShell
Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations.
AprilNEA/OpenLogi
Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.
getpaseo/paseo
Forms a two-agent committee with contrasting profiles to analyze a stuck problem in parallel, reconcile their views and return a consensus plan without editing files.
dotnet/maui
Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.
malloydata/publisher
Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.
malloydata/publisher
Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…
malloydata/publisher
Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.
malloydata/publisher
Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.
malloydata/publisher
Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.
malloydata/publisher
Decide whether ONE answer matches its golden, and say whether you believe the golden.
Diagnose why a scored answer failed and who owns the fix, then cluster the failures by shared root cause. Eval Diagnose is an agent skill from malloydata/publisher. Diagnose why a scored answer failed and who owns the fix, then cluster the failures by shared root cause.
Eval Diagnose fits situations like: tasks that involve Root cause analysis.
Run `npx skills add malloydata/publisher --skill eval-diagnose -a claude-code`. Or copy the skill folder (skills/eval-diagnose in malloydata/publisher) into .claude/skills/eval-diagnose in your project. Claude Code loads it when a task matches its description.
Run `npx skills add malloydata/publisher --skill eval-diagnose -a codex`. Or copy the skill folder (skills/eval-diagnose in malloydata/publisher) into .agents/skills/eval-diagnose in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-diagnose -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-diagnose, .gemini/skills/eval-diagnose, .github/skills/eval-diagnose and .opencode/skills/eval-diagnose in your project.
Going by SKILL.md and its folder, Eval Diagnose needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Eval Diagnose is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.9k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Diagnose: Diagnosing Superpowers Sessions (obra/superpowers, 296k stars), Code Design Rationale Investigator (cursor/plugins, 10k stars), Pester Failure Analysis (PowerShell/PowerShell, 56k stars) and OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 8, 2026.
Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.