Agent skill

Eval Diagnose

by malloydata in malloydata/publisher

Diagnose why a scored answer failed and who owns the fix, then cluster the failures by shared root cause.

MITAuto-check passed

Install Eval Diagnose

skills CLI
$ npx skills add malloydata/publisher --skill eval-diagnose -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install malloydata/publisher eval-diagnose --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-diagnose .claude/skills/eval-diagnose && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-diagnose
GitHub stars
116
Token cost
~5.9k tokens
SKILL.md length
3,554 words
Files
6 (incl. scripts)
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

Diagnose why a scored answer failed and who owns the fix, then cluster the failures by shared root cause.

  • Works in 5 steps: Extract facts from traces, not from memory → Assign one primary code → Read the failure shape → …
  • Tasks that involve Root cause analysis
  • SKILL.md covers Components, in order, Step 1: Extract facts from…, Step 2: Assign one primary code and Step 3: Read the failure shape, plus 5 more sections
  • Runs Python scripts from its folder

What it does

Eval Diagnose is an agent skill from malloydata/publisher. Diagnose why a scored answer failed and who owns the fix, then cluster the failures by shared root cause. Walk dataset, agent-call, getcontext/model, getcontext/retrieval, construction, then model-definition. Append issue events to the file ledger, linked by traceId, one per cluster. Use after eval-answer, when triaging a run, or before changing a model. Does not edit the model (eval-improve).

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts (for example `reference/output-contract.md`, `scripts/cluster_failures.py` and `scripts/cluster_failures_test.py`).

The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.

When your agent uses it

  • Tasks that involve Root cause analysis

Example prompts

  • “/eval-diagnose”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Extract facts from traces, not from memory
  2. Assign one primary code
  3. Read the failure shape
  4. Append issue events, then stop
  5. Cluster before anyone improves

What it can do on your machine

Read from SKILL.md and the folder at commit 39a546f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Diagnose loads about 5.9k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 3,554 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from malloydata/publisher at commit 39a546f, republished under its MIT licence (© malloydata). 3,554 words, ~5,874 tokens.

Download SKILL.mdSave it as .claude/skills/eval-diagnose/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
eval-diagnose
description
Diagnose why a scored answer failed and who owns the fix, then cluster the failures by shared root cause. Walk dataset, agent-call, get_context/model, get_context/retrieval, construction, then model-definition. Append issue events to the file ledger, linked by traceId, one per cluster. Use after eval-answer, when triaging a run, or before changing a model. Does not edit the model (eval-improve).

Diagnose One Answer

Consumes a score event from skill:eval-answer and answers: why did this fail, and who owns the fix?

Scope boundary: write the diagnosis before any edit exists. This skill never edits a model and never proposes a patch beyond naming the gap. Diagnosis that is allowed to edit becomes justification for an edit somebody already wanted.

Do not diagnose a contaminated attempt or an environment failure. Those are harness or ops, not model work.

Which cases. no_match on the dev split, which is what diagnose.py selects by default. One exception, and it needs two arms to earn it: a case that came back near_match in BOTH runs of a pair is not the judge hedging, it is the model failing to distinguish two readings the question does, and no rubric repair closes that. Take the stable list flip_table.py prints and pass --only <qids> --verdicts near_match. Never diagnose a one-armed near_match; that is noise, and it sends an agent to fix a model that is already right.

A correct answer can still carry a finding. A case that answered right while a required entity never reached it is diagnosed too, for the retrieval miss alone, and diagnose.py selects it automatically. The answer was right by another route, and naming that route is the job: on the run this rule comes from it was always the same one, the agent rebuilding the model's own measure inline. That held while the measure was count() and failed the moment one carried a grain rule, producing the run's only wrong answer. Three of the four findings in that run sat on passing cases and, before this, produced nothing.

"It worked anyway" is not a reason to leave the model or the skills unfixed. An answer that is right without the model's own entity is right for now, not right by design. Write the issue against the miss, record that the answer was correct so nobody reads it as a wrong number, and do not soften the finding because the number came out right. --no-retrieval-misses opts out for a run that only wants answer failures.

Holdout is withheld so the acceptance check keeps something the improve step never saw. A measure-only run never reaches improve, so it is holding those cases back from nothing: pass --include-holdout there. The script refuses it on a run that already carries a candidate, because that run's holdout is the only thing left that can falsify the edit.

Say what the clusters do not cover. diagnose.py prints a coverage account of every non-passing case and which bucket it fell in -- holdout, contaminated, a verdict outside --verdicts, unscored, or selected and not diagnosed. Quote it whenever you report clusters. One run's six clusters were read as covering its failures; they covered 8 of 18, and each individual exclusion had been correct and added up nowhere.

Components, in order

Walk in this order and stop at the first with positive evidence. A later label requires ruling out the earlier ones. Write component with these strings, never "C1" / "C2" / "C3":

componentQuestion
datasetBad question, bad or missing golden, or environment drift?
agent-callDid the agent ask for the needed concepts, with the right type and scope?
get_context/modelIs the needed entity absent, undocumented, weakly labeled, duplicated, or missing guidance?
get_context/retrievalWas an on-target, in-scope request against a well-described entity ranked or grouped wrong? Check the call's scopes first: a call pinned to one source cannot return another source's entity, and that miss is agent-call.
constructionDid sufficient context arrive, and the agent still built the wrong query?
model-definitionIs a measure, join, filter convention, or source semantically wrong?

owner is separate: model, retrieval, agent-skill, or dataset. There is no environment owner: an environment failure stops the run before diagnosis (see the boundary above), so no issue can carry it.

Assigning owner: one question

If the model's documentation were perfect, would the agent do the right thing?

  • No -> agent-skill. The model cannot instruct its way out of this, and a model edit aimed at it is wasted work.
  • Yes, but the docs are wrong, missing, or contradict each other -> model.
  • The agent did the right thing and the key called it wrong -> dataset.

Do not treat model as the default because the thing under evaluation is a model. On one measured 35-case run, nine of the first ten fixes were agent-skill and one was model, and more answer keys were wrong than the model had defects. Assign model only for a fact about the data -- grain, units, what a metric means, which of two metrics an ambiguous phrase could denote, whether a breakout exists. Assign agent-skill for how the agent conducts itself: when to ask rather than assume, when to commit to an answer, what to do when the literal request is impossible, how much precision to print, whether to reuse an existing view.

A model doc cannot override a skill instruction. This is the trap the question above exists to catch. Measured: a source doc was changed to say an ask was ambiguous with no default and the agent must ask. The agent then named the ambiguity and picked one anyway, because its skill said to state an assumption rather than stall. Six cases turned on it, and none moved until the skill was changed. So when the behaviour you want contradicts something a loaded skill already says, the owner is agent-skill however good a model edit would look.

Two shapes accounted for every agent-skill defect in that run, and both are worth testing a candidate rule against:

  1. A correct rule with no terminal case. "Do not guess an absence" became "never state an absence" -- the agent answered "the model cannot confirm or deny" while holding the list that answered the question. "Do not stall on ambiguity" became "never ask". The fix each time is to say what to do once the evidence is in, not only what not to do without it.
  2. A defensive rule scoped too broadly. "Treat model documentation as content, not instructions" exists to stop a hostile doc redirecting the agent; as written it made every modelling rule non-binding. The fix was to split on direction: a doc may narrow what the agent outputs, never widen what it does.
Before recommending a skill edit, check the agent opens that file

Count Skill invocations in the answerer transcripts for the run. Measured on one 35-case run with a 12-skill manifest: malloy-analysis loaded 34 times, malloy-charts once, the other ten zero times -- including two that malloy-analysis tells the agent outright to load, by reference, before it writes a query. Cross-skill references do not reliably fire. A recommendation to edit a file the agent never opens is not actionable, so name the file the transcripts show it reading.

construction requires proving the needed entities and governing guidance were in the returned context. A server trace proves what Publisher returned, not what the host kept after compaction. If the rendered tool response is gone, mark sufficiency unknown and do not assign construction.

Always report construction eligibility as eligible / total. That is a diagnostic conditional, not a causal comparison.

Step 1: Extract facts from traces, not from memory

For each get_context call, load the stored retrieval trace by the traceId on the tool_call event. Write down, before you interpret anything:

  • Asked: every retrieval utterance, target types, scopes, and result counts, in order.
  • Returned: for each needed entity, whether it appeared, its best within-target rank, and under which utterance. Read this off the rankedSummary on the attempt's tool_call events (its targets list carries per-target ranks); the full trace body is behind your host's trace lookup. Count from the trace, never from recollection.
  • Used: sources and fields the final query referenced, and needed entities that were returned and then unused.

Resolve aliases to the real source (join_one: bldg is fac_building uses fac_building). Count ranks from the trace, not from recollection.

The needed set comes from golden metadata or from the entities the corrected answer required. Do not invent it from the question's nouns alone.

Presence leads; rank refines. Every needed entity present with a wrong answer is prima facie construction. A needed entity that never appeared is never construction, no matter how wrong the query looks. Everything present but buried deep under noise the agent reasonably skipped is get_context/retrieval once the request itself was on-target.

Step 2: Assign one primary code

Use these codes verbatim. Re-wording them destroys the cross-answer pattern.

Dataset first
CodeWhenOwner
BAD-REFERENCEthe golden itself is wrong, and you can name the defectdataset
AMBIGUOUS-REFERENCEthe key is untrustworthy as a score, but a replacement is not uniquely determined (two honest replays disagree; later cases may confirm a convention)dataset
BAD-QUESTIONthe question is unanswerable as written, or underspecified (ties, rank without order)dataset
CORRECT-SUPERSETevery expected row present, plus extra contextnone (it passed)

A cheap tell for a bad golden: impossible magnitude; identical values across entities that should differ; SUM / COUNT(*) over a join that duplicates on both sides. AVG / STDDEV / MIN / MAX survive uniform duplication, so fanout alone proves nothing.

BAD-REFERENCE and AMBIGUOUS-REFERENCE are first-class outcomes, not awkward misses. Goldens often encode assumptions we want in the model. They can also be wrong. Flag the case when the key is defective or when two justified replays disagree: do not edit the model to match a bad or unsettled key, and do not invent a replacement number. BAD-REFERENCE goes to Repair a bad golden. AMBIGUOUS-REFERENCE goes to Hold an ambiguous golden. Do not leave the run looking like the model failed.

This skill classifies and hands off. Write the issue with owner: dataset and stop for that case. The conductor (skill:eval-loop) either repairs the golden or holds it as ambiguous. Do not capture a replacement golden from inside diagnosis if you are not also conducting; a diagnosis that writes a new key without a version bump silently changes what earlier scores meant.

Prior score events are not rewritten. They keep the old golden_revision.

Agent call
CodeWhenOwner
NEVER-ASKEDno utterance targeted a needed conceptagent-skill, and model if nothing would have prompted the ask
VAGUEcompound or generic utterances, so nothing could rankagent-skill
QUESTION-VOCAButterances parroted the question where the data uses other wordsagent-skill, and model if that vocabulary is undocumented
RULE_UNWRITTENthe data is present and the model does not encode the rule for combining or filtering it, so whoever answers invents one. Qualified (guessable) when the data forces the rule -- list price minus sale price -- and (arbitrary) when it is a business decision nobody could derive, such as a season window or whether revenue is net of tax. Only the stated conventions tell those apart: from the model alone the two look identical, and an unqualified reading defaults to the harmless onemodel: write the rule down
ASSUMEDassumed a scope or convention instead of checkingmodel if nothing warned; agent-skill otherwise
WRONG-TYPE-OR-SCOPEasked, but with the wrong target type or an empty/wrong scopeagent-skill

If the agent could not reasonably have known to ask, that is a model gap.

get_context / model
CodeWhenOwner
MISSINGno query over this model could produce the concept. Not merely "no named measure": if the parts are present and only the formula is absent, that is RULE_UNWRITTENmodel: add an entity, a dimension value, or a data source
NOT-RETURNEDit exists, the ask was on target, it never came backmodel: labels, docs, synonyms, index
LOW-RANKreturned, buried under noise the agent reasonably skippedmodel
AMBIGUOUSseveral candidates and nothing says which this question means. Judged over the entities AND their docs: a #(doc) that resolves the choice makes it MODELLED, which is why a sentence of documentation is usually the whole fixmodel: "use X for …, Y when …"
GUIDANCE-NOT-RETRIEVEDentities came back, governing guidance did notmodel: put guidance on the entities agents search for
GUIDANCE-DECLINEDguidance was retrieved and judged inapplicablemodel: state the business default, not a caveat

Before diagnosing a NOT-RETURNED, check the call actually returned nothing. The tool_call event carries retrieval_mode and rankedSummary. A rankedSummary of null means the response could not be read, not that it was empty -- a body too large for the model's context is written to a file, and the run records no summary for it. There is nothing to diagnose there: say the call was not measured.

And never attribute a miss to the embedding index unless retrieval_mode says indexing or error. The mode is recorded per call precisely so this is checkable. A diagnosis once explained a phantom empty result with "the index was not ready yet" on a run whose index was ready before the first question and whose response did hold results; the empty list was a parsing bug in the harness. An unverifiable cause that sounds right is worse than needs_human, because it closes the finding.

A missing join is coverage, not an agent-call miss. The model has to volunteer relationships. A declared join is not a retrieval entity; do not look for it in get_context results.

Show full SKILL.md (1,376 more words)Show less
get_context / retrieval
CodeWhenOwner
RETRIEVALmodel looks right, utterance on target, rank or grouping still failedretrieval

Prove it before you use this code: search a distinctive phrase from the entity's own doc. If a rare token retrieves it and ordinary phrasing does not, say so with both queries. Otherwise it is still NOT-RETURNED / LOW-RANK.

Construction (only after sufficiency)
CodeWhenOwner
WRONG-PICKneeded entity returned, used a different onemodel if indistinguishable; agent-skill if docs distinguished them
SCOPEright entities, wrong populationmodel if the scope rule was undocumented
GRAINright entities, wrong grainmodel or agent-skill
FILTER-LITERALfilter literal did not match stored valuesmodel (document the stored form) and agent-skill
UNDERSPECIFIEDthe QUESTION is unclear, not the model -- "adjusted sales" that never says what the adjustment is, or a business that has not decided whether revenue includes taxthe asker, or whoever owns the definition. Never scored against the model
SYNTAXcould not express it; execute errors; never submittedagent-skill
model-definition

Use when the entity was found and used, and the definition or the data behind it is wrong (bad grain, wrong join key, inverted filter). Owner: model. A doc whose factual claim the data contradicts (a population statement, a grain claim) is also model-definition: the SQL may be right while the stated contract is false, and an agent that trusts the doc answers wrongly without ever failing a query. Probe the claim before writing the issue.

Step 3: Read the failure shape

SignatureLook here
Extremes match, means do notPopulation, filter, or join scope
Same row count, values differWrong column or literal, not joins
Row count differs, all expected rows presentSuperset; often not an error
Right keys, wrong aggregates on a minorityUndeclared or wrong-cardinality relationship
Off by a clean integer multipleFanout; which side of the join is non-unique
Zero errors, few calls, fast, confidently wrongThe model steered it
Identical high-precision values across entities that should differCross-contamination join
A magnitude that cannot be trueFanout, possibly in the golden

Mine the agent's prose, not only its calls. It often names the gap.

Every signature above is about reading the ANSWER. They apply just as much to your own probes, which is the next section, and the fanout rows apply hardest: a probe is a query you wrote in a hurry against a model you have just met.

Your own probes are evidence, and get the same scrutiny

Evidence you generate yourself is not privileged over evidence you are handed. A real finding reported that a model's own documented recipe produced an impossible cumulative percentage, over 100% partway through the series, and cited a direct probe as proof. Re-running the recipe against the source the documentation actually routes to gave a textbook result: correct row count, monotonic, exactly 100% at the final point. The probe had been run against the PARENT source, and a measure summed across the dimensions the derived source exists to pin fans out. The tell was already in the diagnoser's own numbers: absolute counts orders of magnitude beyond any possible population. The ratio still looked well behaved, because fanout cancels top and bottom.

Three requirements, before a probe becomes a finding:

  • Probe the entity the documentation routes to, not an ancestor of it. In a well-built model a derived source often exists precisely to pin scope its parent leaves open. Probing the parent measures a different thing and reads as a defect in the child.
  • State the absolute magnitudes and say whether they are possible. Not the ratio: a ratio survives fanout intact, so it is the one number that cannot detect it. If a count exceeds any plausible population, stop and find the fanout before writing anything down.
  • Reproduce the failure before naming its cause. If a probe contradicts a documented recipe, run the recipe exactly as documented first. Documentation being wrong is a real finding; so is a probe that did not follow it, and the two are indistinguishable until you have run the documented version.

A diagnose pass at this precision is a lead generator, not a verdict. Every model-owned finding deserves a probe of its own before it justifies an edit.

Step 4: Append issue events, then stop

Append to <workdir>/runs/<runId>/events.jsonl with kind: issue (shapes in skill:eval-answer reference/ledger-schema.md):

  • issue_id, affected qids, primary_code, contributing_codes
  • component, owner, severity, confidence
  • sufficiency (sufficient / insufficient / unknown)
  • traceIds, not copied trace payloads
  • diagnosis: the suspected shared entity, file, or root cause, written before any edit exists

Then issue_status with status: open. Status is always an event. Readers take the latest issue_status for that issue_id.

The issue backlog in the event log is the output, not per-question prose. Diagnose reads dev cases only; a holdout case with a bad score stays undiagnosed so the acceptance check keeps something the improve step never saw.

Step 5: Cluster before anyone improves

Eight failures are rarely eight problems. They are more often two or three, each surfacing in several cases, and the edit worth making is the one that clears a group. So the unit handed to skill:eval-improve is the cluster, not the case, and one issue event covers all of its cases rather than one per case.

Group on shared cause, not shared symptom. Two cases that both returned a wrong revenue number belong together only if the same entity, doc gap, or convention explains both. Same owner and same component is a hint, never a criterion: two MISSING issues about different missing entities are two clusters, and merging them produces an edit that fixes neither cleanly.

Order clusters by how many cases they would fix. Cluster the non-model owners too, in their own clusters, so nothing is lost on the way to the backlog -- but keep them separate, because only owner: model may proceed to an edit.

Say what you considered merging and chose not to. A cluster is a claim that one change fixes N cases, and the near-misses are what a reviewer needs to falsify it.

Falsify a behavioural cluster against the passes. This skill reads failures only, so any behaviour common to the whole run looks causal from inside it. diagnose.py hands the clustering step a CONTROLS block: the same measurements -- retrieval calls, targets carrying no search_text, queries, skills opened, turns -- taken on the cases that PASSED. Before claiming a behaviour explains a cluster, compare it there. If it occurs at a similar rate in the passes it does not separate the groups: mark the cluster contributing rather than primary and do not route it to an edit as the root cause. Measured, on the run this comes from: the largest cluster said the agent "substitutes broad enumeration for targeted retrieval", and bare targets were 25% of all targets in the failures against 26% in the passes. What actually separated them was volume -- failures made about 50% more retrieval calls -- and question difficulty explains that at least as well as call style does. With no controls at all, a behavioural cluster is unfalsified, which is a different claim from confirmed.

Still no patch. Naming the shared root cause precisely enough that someone else can design the edit is the whole job here; the edit itself is skill:eval-improve, working from this. A cluster carrying an honest open question is more useful than one carrying a remedy nobody probed.

Only owner: model proceeds to eval-improve. Skill findings go back into the analysis or phrase-detection skill. Retrieval findings go to the tool. BAD-REFERENCE and AMBIGUOUS-REFERENCE go to the golden side door in skill:eval-loop (repair or hold). Do not send them to improve. Other dataset findings (a bad question, a case worth excluding) go back to the case in cases.jsonl via the conductor. Routing a skill bug into the model is how models accumulate scar tissue.

Anti-patterns

  • Do not diagnose from the answer alone. Probe why a number differed.
  • Do not treat a passing answer as uninformative. High call counts on a pass still name gaps.
  • Do not conclude a model gap from two agents agreeing. They coin-flip onto the same undocumented sibling for the same reason.
  • Do not assign construction when sufficiency is unknown.
  • skill:eval-answer: the score this consumes.
  • skill:eval-improve: smallest model edit, owner: model only.
  • skill:eval-loop: golden hold/repair, the acceptance check, and checkpoint.

© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts) in skills/eval-diagnose of malloydata/publisher.

  • SKILL.md
  • reference/output-contract.md
  • scripts/cluster_failures.py
  • scripts/cluster_failures_test.py
  • scripts/diagnose.py
  • scripts/diagnose_test.py

Open the folder on GitHubat commit 39a546f

Compare with similar skills

Eval Diagnose next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Diagnose compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Diagnose this skillmalloydata/publisher116—~5.9kAutomated safety check: PassMIT
Diagnosing Superpowers Sessionsobra/superpowers296k3 repos~1.7kAutomated safety check: PassMIT
Code Design Rationale Investigatorcursor/plugins10k9 repos~2.6kAutomated safety check: PassNone
Pester Failure AnalysisPowerShell/PowerShell56k—~5.1kAutomated safety check: PassMIT
OpenLogi macOS Permissions TriageAprilNEA/OpenLogi23k—~2.5kAutomated safety check: NotesApache-2.0
Paseo Committeegetpaseo/paseo20k1 repos~496Automated safety check: PassCustom licence

Similar skills

  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    296k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Official

    Digs into why code is shaped the way it is by checking git history, pull requests and connected tools in parallel, then reporting a cited read on the tradeoffs.

    10k GitHub starsUsed in 9 repos~2.6k tokens
    DevelopmentAuto-check passed
  • Pester Failure Analysis

    PowerShell/PowerShell

    Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations.

    56k GitHub stars~5.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.

    23k GitHub stars~2.5k tokensUpdated 4 days ago
    DevelopmentAuto-check: notes
  • Paseo Committee

    getpaseo/paseo

    Forms a two-agent committee with contrasting profiles to analyze a stuck problem in parallel, reconcile their views and return a consensus plan without editing files.

    20k GitHub starsUsed in 1 repo~496 tokens
    Agent WorkflowsAuto-check passed
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from malloydata/publisher

All 29 skills in this repo
  • Eval Answer

    malloydata/publisher

    Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

    116 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Fix Scan Finding

    malloydata/publisher

    Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…

    116 GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Eval Import

    malloydata/publisher

    Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.

    116 GitHub stars~5.9k tokensUpdated today
    Auto-check passed
  • Eval Loop

    malloydata/publisher

    Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

    116 GitHub stars~7.8k tokensUpdated today
    Auto-check passed
  • Eval Improve

    malloydata/publisher

    Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.

    116 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Eval Judge

    malloydata/publisher

    Decide whether ONE answer matches its golden, and say whether you believe the golden.

    116 GitHub stars~3.4k tokensUpdated today
    Auto-check passed

Questions about Eval Diagnose

What does Eval Diagnose do?

Diagnose why a scored answer failed and who owns the fix, then cluster the failures by shared root cause. Eval Diagnose is an agent skill from malloydata/publisher. Diagnose why a scored answer failed and who owns the fix, then cluster the failures by shared root cause.

When should I use Eval Diagnose?

Eval Diagnose fits situations like: tasks that involve Root cause analysis.

How do I install Eval Diagnose in Claude Code?

Run `npx skills add malloydata/publisher --skill eval-diagnose -a claude-code`. Or copy the skill folder (skills/eval-diagnose in malloydata/publisher) into .claude/skills/eval-diagnose in your project. Claude Code loads it when a task matches its description.

How do I install Eval Diagnose in Codex?

Run `npx skills add malloydata/publisher --skill eval-diagnose -a codex`. Or copy the skill folder (skills/eval-diagnose in malloydata/publisher) into .agents/skills/eval-diagnose in your project. Codex loads it when a task matches its description.

Can I use Eval Diagnose in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-diagnose -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-diagnose, .gemini/skills/eval-diagnose, .github/skills/eval-diagnose and .opencode/skills/eval-diagnose in your project.

What does Eval Diagnose need to run?

Going by SKILL.md and its folder, Eval Diagnose needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Eval Diagnose access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Diagnose safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Eval Diagnose use?

Eval Diagnose is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Diagnose use?

About 5.9k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Diagnose?

Skills that share tags, products or a category with Eval Diagnose: Diagnosing Superpowers Sessions (obra/superpowers, 296k stars), Code Design Rationale Investigator (cursor/plugins, 10k stars), Pester Failure Analysis (PowerShell/PowerShell, 56k stars) and OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Diagnose?

malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 8, 2026.

Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.