Data Table Manager
n8n-io/n8n
Load before calling data-tables or parse-file. An agent skill from n8n-io/n8n.
Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.
$ npx skills add malloydata/publisher --skill eval-import -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install malloydata/publisher eval-import --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-import .claude/skills/eval-import && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-import" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-import into .claude/skills/eval-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-import", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/malloydata/publisher/tree/main/skills/eval-importType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add malloydata/publisher --skill eval-import -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install malloydata/publisher eval-import --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-import .agents/skills/eval-import && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-import" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-import into .agents/skills/eval-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-import", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-import -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install malloydata/publisher eval-import --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-import .cursor/skills/eval-import && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-import" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-import into .cursor/skills/eval-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-import", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/malloydata/publisher.git --path skills/eval-import--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add malloydata/publisher --skill eval-import -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install malloydata/publisher eval-import --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-import .gemini/skills/eval-import && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-import" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-import into .gemini/skills/eval-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-import", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install malloydata/publisher eval-importInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add malloydata/publisher --skill eval-import -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-import .github/skills/eval-import && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-import" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-import into .github/skills/eval-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-import", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-import -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install malloydata/publisher eval-import --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-import .opencode/skills/eval-import && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-import" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-import into .opencode/skills/eval-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-import", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-importTurn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.
Eval Import is an agent skill from malloydata/publisher. Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs. Classify each item by what came WITH the question (a query, a number, prose criteria, or nothing), write cases.jsonl and set.json per reference/case-format.md, keep the file as it arrived, and seal each question so a later edit is detectable. Never marks a value-bearing golden verified and never invents a question. Use when…
Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `reference/case-format.md`, `scripts/import_cases.py` and `scripts/import_cases_test.py`).
It sits in Documents & Office, covering Excel spreadsheets and CSV and tabular files. The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 39a546f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Import loads about 5.9k tokens when it runs. Until then it costs about 161 tokens; SKILL.md has 3,722 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from malloydata/publisher at commit 39a546f, republished under its MIT licence (© malloydata). 3,722 words, ~5,853 tokens.
.claude/skills/eval-import/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Questions arrive in whatever shape their author had. This turns them into
evals/<set>/, ready for a run, and records what each key is actually worth.
Scope boundary: cases only. This never answers a question, scores an
answer, or edits a model. It is skill:eval-loop's scrape step, expanded,
and it is where the set's honesty is decided.
The file shapes are defined once, in reference/ledger-schema.md of
skill:eval-answer. Read them there; this skill does not restate them.
reference/case-format.md covers what is specific to an import: which fields
an arriving item maps to, and the four physical formats.
Two separate rules, and "package" blurs them:
So the set is a SIBLING of the package under test, in its repo. A repo with
ecommerce/ under test gets evals/ecommerce-questions/, beside it.
Copy it verbatim into evals/<set>/as-received/, and never edit that copy.
Not the encoding, not the line endings, not an obvious typo.
It is there so a later disagreement is about the key and not about whether
somebody transcribed the question correctly. That disagreement is guaranteed:
skill:eval-loop's reference/auditing-an-answer-key.md is about what happens
when it arrives, and its first check is the question against this file.
If the source cannot be committed, because it carries names, addresses,
internal URLs, or row values somebody pasted into a thread, say so. Record in
set.json where the original lives and a hash of it, note under
sourceNote that the copy is withheld and why, and tell the user that the
byte-for-byte check is now impossible for anyone without the original. Do not
redact silently: a redaction you got wrong becomes the provenance record.
This is the whole job, and getting it wrong is how an unverified number gets scored as a fact.
| what arrived | write | scorable on day one |
|---|---|---|
| the question only | a case with no golden | no. It measures coverage and discoverability, which is a real measurement |
| question + their query + a number | golden.value, their query as canonicalQuery, verifiedBy: authored_query, status: provisional | after step 3 agrees |
| question + a number, no query | golden.value, status: provisional, no verifiedBy | no. Somebody typed a number |
| question + prose criteria | see step 4 | depends on what the criterion is |
Two rules over the whole table:
No golden holding a VALUE imports as verified. verified means two
differently shaped derivations agreed through the truth package, and an import
has performed neither. provisional goldens are unscorable by design:
skill:eval-answer issues verdict: null on them. That is not a gap to work
around.
It is also not a dead end, and it used to read like one. The way out is
verify_goldens.py --promote, which marks a provisional golden verified
once its value re-derives cleanly from the truth package and a second
derivation exists. It is the only thing anywhere that writes golden.status.
Tell the user that command when you hand over a set of provisionals, because
until it is run the set scores nothing. A set that arrives with 60 typed numbers genuinely has 60 unverified
numbers, and the number it can honestly produce on day one is a coverage score.
The exception is a golden that holds no value at all, and there are two
kinds. Prose criteria (kind: criteria, step 4) have nothing to re-derive:
the author's words ARE the key. So does a case whose stated answer is that the
data is not there (kind: unanswerable), where the pass is a refusal that
names what is missing and a confident number is the failure. Both import
verified with verifiedBy: authored_criteria and score immediately. The
rule is about numbers, because a number is the thing that can be wrong while
looking right.
Watch for the second one in what arrives. A criterion reading "a refusal that
names the missing data" is not a rubric clause on a normal case, it is the
whole key, and importing it as criteria on a case with a value would make a
refusal fail.
A question with no golden is a case, not a reject. Do not hold questions back until somebody derives keys. A set of bare questions already measures whether the model can express an answer at all, and the answers it produces are what the keys get derived from.
The table above is about what ARRIVED. Sometimes nothing did, nobody knows the right number, and the set still needs keys. You may derive one, and the rules are the ones already stated rather than new ones:
Use your full model access and query until you can defend the number. The
agent that gets tested later holds only get_context and execute_query;
you are not that agent, and the gap is the point.
Record the query as canonicalQuery. A number with no query is somebody's
guess, which is the row above it in the table.
Record the entities you used to get there as expectedEntities.required, in
the kind:source:name form: the measures, dimensions and views the answer
cannot be produced without. reference/case-format.md says to leave this
field empty at import because guessing it invents a retrieval expectation
nobody stated; you are not guessing, you just used them. That is the only
trustworthy producer this field has, and it is what makes retrieval
measurable on a set nobody sent keys for. Where the model offers two
legitimate routes, write a requiredAnyOf group rather than pick one.
verify_goldens.py check 5 audits every id against the model, so a wrong one
is a hard finding rather than a silent retrieval miss.
Derive the list with the model in front of you, not from memory, and not
with grep. The model knows what each field IS; an author does not reliably.
Ask the question of the model: which entities does an answer to this depend
on, and is each a measure or a dimension. Confirm every id with get_context
before writing it. A list written from recollection is the most common way a
set acquires a retrieval expectation nobody can satisfy.
Measured: an agent given the model and one question found alternatives a
hand-written key had missed on all three cases tried -- carrier name against
nickname against code, destination_count against its underlying
destination.airport_count. At roughly $0.13 a question. It is better than a
person working from notes at finding what the model actually offers.
It is worse at deciding what to REQUIRE, and the two rules below are why. On the same three cases it proposed demoting two named measures on the grounds that a raw expression answered identically. One of those equivalences was false; the other was true today and is not a reason to demote anything.
But VERIFY every "this is not strictly required" claim by running both sides. This is the rule that makes the step above safe, and it is not optional.
The same probe got one badly wrong. Told that a named measure is not required
when a raw column plus a plain aggregate answers identically, it wrote that
average_plane_size was "equally answerable by
avg(aircraft.aircraft_models.seats)". Executed, those are 229 and 196.93.
The measure is aircraft.avg(...), an average over distinct aircraft, and the
inline version averages over flights. The claim was false for exactly the
measure whose entire purpose is to carry that distinction -- and it is the
same mistake, on the same field, that produced the only wrong answer in the
run this set came from.
Had that proposal been written into the key, the case would have stopped testing the thing it exists to test.
So: an equivalence is a claim about DATA, and reading a definition is not evidence. Run both expressions and compare the rows.
But agreeing today is not a reason to demote the model's entity, and this
is the part the probe got backwards. airport_count is count() returns the
same 984 as a bare count() right now, and stops doing so the moment anyone
adds a null filter or a deduplication to it. The key would not notice: it
would go on passing while testing nothing. That applies to every measure that
is currently a thin wrapper, which is most of them until the day one is not.
requiredAnyOf is for two routes the MODEL defines -- a carrier's name
and its nickname, destination_count and the destination.airport_count
it wraps. Both are entities, both are retrievable, and an answer using either
has gone through the semantic layer. That is a real choice and the group
records it.
It is NOT for "the model's entity, or a raw expression that happens to
agree". Those are not two routes: one is the route, the other is a bypass
that currently works. Keeping the named entity in required is what makes
the eval measure what the semantic layer is for, and finding it IS the
retrieval behaviour under test -- an agent that searches for the pre-defined
entity inherits its definition and its rules, while one that rebuilds from
raw columns inherits whatever it thought of.
What the equivalence check is FOR is reading a run afterwards. When an answer came out right without the entity, comparing the two says whether it got lucky or got it wrong, and a set whose passes are merely lucky is a set of near-misses.
Then prove each one is retrievable, and record what finds it.
check_findable.py --set <set> --mcp-url <mcp> --publisher <rest> --environment <env> --package <pkg> runs two checks and needs no model call.
With --publisher it reads the COMPILED model and settles whether each field
exists and what KIND it is; then it searches for each entity to confirm the
index can actually deliver it.
This matters more than it looks, because required is what BOTH retrieval
numbers are computed against: whether the agent asked for the entity, and
whether it came back. An id retrieval cannot deliver makes both fiction, and
the case reports a retrieval miss on every run, which reads as a defect in
the model or the agent rather than in the key.
Prefer the compiled model over a grep, and prefer it in both directions.
Measured on one package: dimension:flights:flight_count names a field the
model really has, so grepping the text passes it, and the compiled model says
flight_count is a MEASURE, which no dimension request can return.
Conversely dimension:airports:own_type appears zero times in the .malloy
and the compiled model declares it, because the source exposes it implicitly
from the data; the grep calls it missing and acting on that deletes a good
entity. verify_goldens.py check 5 is the grep, it is free and needs no
server, and it is not the authority.
A pass is a floor, not a verdict on the docs: each entity is searched by its
OWN NAME, the easiest query that could find it. An entity that answers to its
identifier and not to the words a question uses still fails at run time, as
not retrieved.
Writing the list forces the question an author has to answer anyway: what would a reasonable agent search for, and of what type, to find this? If you cannot state that, the agent cannot be expected to guess it.
status: provisional, never verified, for the reason stated above: you
derived it THROUGH the model under test, so a model bug would certify its
own key. verify_goldens.py --promote is still the only way out.
Where two readings are both defensible from the model, do not pick one.
A model that documents two conventions for the same population produces two
honest numbers, and choosing quietly is how a confident wrong key gets
written. Go to Hold an ambiguous golden in skill:eval-loop's
reference/golden-side-door.md: record the competing candidates rather than
a new number.
What no query can settle is not a failure to derive. It is step 4's second kind, and those clauses score on day one.
Do it at import, before any run. It is the cheapest finding in the whole loop.
Execute their query with execute_query against the model under test and
compare with the number they stated. Three outcomes, three different things
learned:
verifiedBy: authored_query. Keep
status: provisional: this proves the number came from that query, and NOT
that the query is right. It ran against the model under test, so a model bug
certifies its own golden, which is the exact circularity the truth package
exists to break. It is still far more than a typed number is worth.provisional. Do not quietly adopt either one: their number may be stale,
the model may have moved, or the query may have always been wrong, and
which of those it is decides who owns the fix.Never repair their query to make it run. A query that does not compile against this model is evidence about this model.
Decide which, per criterion. The test:
Can you write a query whose result settles it?
If yes, it is a value in disguise. "Should be about 4.2M." "The top
category is Denim." "Should come back with twelve rows." There is one right
answer, so the criterion is not the key, it describes one. Write it as
golden.rubric, derive the value, and until it is derived the case is
provisional and unscorable. Grading such a criterion as prose is how an
unverified number passes: the judge reads "about 4.2M", the answer says 4.2M,
and nothing ever checked whether 4.2M is right.
If no, it is a judgment about the shape of the answer, and prose is the
right home for it. "Must break the total out by region." "Must say the window
excludes returns." "Must not use list price." No query settles these, the
judge grades them against the answer, and these cases score on day one. Write
the golden as kind: criteria with no value, the criteria as
golden.rubric, and verifiedBy: authored_criteria. Use golden.mustState
and golden.mustNotUse where the criterion fits those fields exactly, per
reference/case-format.md.
Most arriving criteria are mixed: one sentence naming a number and a shape. Split it. The number half goes provisional; the shape half scores.
Where you cannot decide, mark the case and report it rather than guessing. A
criterion nobody could classify is a question for its author, and
skill:eval-loop's golden side door is where it waits.
Each of these cost a scored case on a real set, and none of them is obvious while you are writing one.
The golden rows are the figures. The rubric's prose is a gloss on them, and
where the two disagree the prose is what is stale. On one set 26 of 29
rubrics quoted a figure that appears nowhere in their own golden rows: the
goldens were re-derived the next morning, the prose was not, and an agent that
computed 747 and 370 -- the exact numbers in the golden JSON -- was failed
against a rubric still saying 615 and 502. Say how to derive the figure, not
what it equalled. verify_goldens.py check 2 reports figures in the accepting
clause that are absent from the rows, but it reads only figures specific enough
to be a quoted result -- a bare three-digit number is invisible to it, and 615
is exactly that. So the check is a help, not a guarantee, and the rule above is
yours to keep. The judge is told the same thing from the other side: where a
rubric and a golden disagree about a figure, it scores against the golden.
Do not assert the model's current behaviour. "contract_terms cannot be
used here at all -- it returns ZERO rows" was true when written and false four
hours later, when a commit unblocked that view. Nothing linked the two, and the
judge then reasoned from the stale claim against an answer that was right.
Co-locating the set with the model makes such drift visible in a diff; it does
not detect it. A rubric that describes a bug is a rubric with an expiry date:
write the requirement, not the defect.
When a question admits two honest populations, accept either and require the answer to name which. "Products in a category" can mean listed-in or primary-category, and on one set the two readings differed by 4,070 against 2,707. Six or more cases turned on it and one rubric had been written to accept both. That one was right. The rest were repaired afterwards, having failed correct answers in the meantime.
If the rubric accepts an alternative, the entity list must too. The two
halves of a key are read by different things: the rubric is prose for the judge,
expectedEntities.required is ids for retrieval scoring. They can disagree
without anything noticing. A rubric saying "either the full carrier name or the
nickname is fine" beside a required list naming only dimension:carriers:name
scores an answer that used the nickname CORRECT and docks it recall in the same
run, and that lost recall then reads as a retrieval failure. Use a
requiredAnyOf group, which is satisfied when any member is delivered.
verify_goldens.py reports the mismatch as a review item, but the rule is
yours: it is a heuristic over prose and cannot catch every phrasing.
A trap note is not a requirement. Notes that arrive beside the questions describe what their author thought was hard, and they mention things the question never asked for. Turning one into a rubric clause invents a requirement the answerer was never given, and it happened: a term appearing nowhere in any question became a clause an answer was marked down for missing. Convert a note only where it constrains the answer to the question as asked.
The question is the stimulus. It is what a human asked, and it is not yours to
tidy. Do not fix a typo, normalize casing, expand an abbreviation, or sharpen a
vague phrase. Vagueness is often the thing the case tests, and narrowing a
question to match what an answerer keeps doing deletes the test and hands the
model a pass. reference/auditing-an-answer-key.md has the audit where that
happened to four questions.
At conversion, stamp questionSha: the SHA-256 of the exact question text, as
written. scripts/import_cases.py --stamp does it, and refuses to overwrite a
stamp that already exists.
The seal is not a derivation from the arriving file, which is why it works for
an email thread as well as a CSV. It records the decision you made at
conversion about what the question is. From then on, verify_goldens.py
compares the stamp against the question in cases.jsonl on every audit, and a
mismatch means somebody edited a question after it was imported.
A question that genuinely has to change gets a new qid, not a new stamp.
A changed question is a different stimulus, and scores on the old wording must
not roll into the new one.
Every case gets split: dev or holdout. Diagnose and improve read dev only;
the acceptance check runs both, and a set that is all dev cannot defend an
accept. skill:eval-loop owns the rule; import is where it is frozen, because
a split chosen after the first failures is not a holdout.
Prefer variety over volume. Cases differing in grain, source, filter shape and phrasing are what move a measurement.
python3 scripts/import_cases.py --set evals/<set> --stampIt checks the required fields, unique qids, a split on every case, a golden
status from the allowed statuses, that nothing claims verified without the
evidence for it, and that no question has drifted from its stamp. Exit 0 clean,
1 with a finding, 2 on a usage error.
Then tell the user, in these terms, what arrived:
47 cases from 50 lines
12 scorable now
29 provisional (21 with their query, 5 a number alone, 3 nothing to compare yet)
6 no golden (question only)Say what turns the 29 into scorable cases, in the same breath: a truth package
(init_truth_package.py), then verify_goldens.py --set <set> --publisher <truth> --promote. A reader told only the counts has no way to know the set is
not simply broken, and the provisional bucket is the one that looks like
progress and is not.
The three provisional buckets are three different amounts of work, which is why they are counted apart. "With their query" needs a re-run. "A number alone" needs a derivation. "Nothing to compare yet" is a case whose criteria describe a key nobody has derived, and it is the one most easily mistaken for progress: measured on a real markdown thread, all three of its asks landed there.
Say the scorable count out loud when you report it, not just the case count. "47 cases" reads like a 47-case measurement, and 12 is the number a first run can actually score. Read the two numbers on the first line against each other too: 47 from 50 means three lines did not parse, and the findings say which.
verified;evals/ inside the served package tree;skill:eval-answer.skill:eval-loop: the conductor. Import is its scrape step; its
reference/auditing-an-answer-key.md is the procedure for a key you doubt,
and its golden side door is where an unclassifiable item waits.skill:eval-answer: reference/ledger-schema.md defines every file and
field this skill writes.© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts) in skills/eval-import of malloydata/publisher.
Open the folder on GitHubat commit 39a546f
Eval Import next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Import this skillmalloydata/publisher | 116 | — | ~5.9k | Automated safety check: Pass | MIT | |
| Data Table Managern8n-io/n8n | 207k | — | ~2.3k | Automated safety check: Pass | Custom licence | |
| Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences | 274 | 2 repos | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| Markitshift-labs-ai/markit | 1.3k | — | ~299 | Automated safety check: Pass | MIT | |
| Convert Fileduckdb/duckdb-skills | 600 | 1 repos | ~720 | Automated safety check: Notes | MIT | |
| Research Integrity Auditxuzhougeng/wisp-science | 1k | — | ~2.6k | Automated safety check: Pass | AGPL-3.0 |
n8n-io/n8n
Load before calling data-tables or parse-file. An agent skill from n8n-io/n8n.
aws-samples/amazon-bedrock-agents-healthcare-lifesciences
Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.
shift-labs-ai/markit
Convert files and URLs to Markdown. An agent skill from shift-labs-ai/markit.
duckdb/duckdb-skills
Convert any data file to another format: CSV, Parquet, JSON, Excel, GeoJSON, and more.
xuzhougeng/wisp-science
学术审查 / research-integrity screening of a manuscript's figures and reported numbers.
eclipse-rdf4j/rdf4j
Parse JMH result text by finding the first header line that starts with Benchmark and contains Mode and Score, build a structured table for all columns/rows, compare overlapping benchmarks across 2+…
malloydata/publisher
Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.
malloydata/publisher
Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…
malloydata/publisher
Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.
malloydata/publisher
Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.
malloydata/publisher
Decide whether ONE answer matches its golden, and say whether you believe the golden.
malloydata/publisher
Combine validated Malloy queries into a notebook report. An agent skill from malloydata/publisher.
Categories
Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs. Eval Import is an agent skill from malloydata/publisher. Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.
Eval Import fits situations like: questions arrive from outside and need to become a set; before the first run against a set nobody here authored.
Run `npx skills add malloydata/publisher --skill eval-import -a claude-code`. Or copy the skill folder (skills/eval-import in malloydata/publisher) into .claude/skills/eval-import in your project. Claude Code loads it when a task matches its description.
Run `npx skills add malloydata/publisher --skill eval-import -a codex`. Or copy the skill folder (skills/eval-import in malloydata/publisher) into .agents/skills/eval-import in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-import -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-import, .gemini/skills/eval-import, .github/skills/eval-import and .opencode/skills/eval-import in your project.
Going by SKILL.md and its folder, Eval Import needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Eval Import is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.9k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Import: Data Table Manager (n8n-io/n8n, 207k stars), Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars), Markit (shift-labs-ai/markit, 1.3k stars) and Convert File (duckdb/duckdb-skills, 600 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 8, 2026.
Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.