Agent skill

Eval Import

by malloydata in malloydata/publisher

Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.

MITAuto-check passedDocuments & Office

Install Eval Import

skills CLI
$ npx skills add malloydata/publisher --skill eval-import -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install malloydata/publisher eval-import --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-import .claude/skills/eval-import && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-import
GitHub stars
116
Token cost
~5.9k tokens
SKILL.md length
3,722 words
Files
4 (incl. scripts)
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.

  • Works in 7 steps: keep the file as it arrived → classify each item by what arrived WITH… → run their query, if they gave one → …
  • Questions arrive from outside and need to become a set
  • SKILL.md covers Where the set goes, Step 1: keep the file as it…, Step 2: classify each item by… and Step 3: run their query, if…, plus 6 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Eval Import is an agent skill from malloydata/publisher. Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs. Classify each item by what came WITH the question (a query, a number, prose criteria, or nothing), write cases.jsonl and set.json per reference/case-format.md, keep the file as it arrived, and seal each question so a later edit is detectable. Never marks a value-bearing golden verified and never invents a question. Use when…

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `reference/case-format.md`, `scripts/import_cases.py` and `scripts/import_cases_test.py`).

It sits in Documents & Office, covering Excel spreadsheets and CSV and tabular files. The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.

When your agent uses it

  • Questions arrive from outside and need to become a set
  • Before the first run against a set nobody here authored

Example prompts

  • “/eval-import”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. keep the file as it arrived
  2. classify each item by what arrived WITH the question
  3. run their query, if they gave one
  4. prose criteria are two different things
  5. never change a question, and seal it
  6. freeze the split
  7. validate, and report what the set is worth

What it can do on your machine

Read from SKILL.md and the folder at commit 39a546f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Import loads about 5.9k tokens when it runs. Until then it costs about 161 tokens; SKILL.md has 3,722 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~161
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from malloydata/publisher at commit 39a546f, republished under its MIT licence (© malloydata). 3,722 words, ~5,853 tokens.

Download SKILL.mdSave it as .claude/skills/eval-import/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
eval-import
description
Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs. Classify each item by what came WITH the question (a query, a number, prose criteria, or nothing), write cases.jsonl and set.json per reference/case-format.md, keep the file as it arrived, and seal each question so a later edit is detectable. Never marks a value-bearing golden verified and never invents a question. Use when questions arrive from outside and need to become a set, or before the first run against a set nobody here authored.

Import questions into an eval set

Questions arrive in whatever shape their author had. This turns them into evals/<set>/, ready for a run, and records what each key is actually worth.

Scope boundary: cases only. This never answers a question, scores an answer, or edits a model. It is skill:eval-loop's scrape step, expanded, and it is where the set's honesty is decided.

The file shapes are defined once, in reference/ledger-schema.md of skill:eval-answer. Read them there; this skill does not restate them. reference/case-format.md covers what is specific to an import: which fields an arriving item maps to, and the four physical formats.

Where the set goes

Two separate rules, and "package" blurs them:

  • The same git repository as the model it evaluates, so one commit pins the model and the ledger together. That is what a checkpoint is.
  • Never inside the directory tree the answerer's package serves. The answerer can read files in the served tree, and gold there is a contamination path.

So the set is a SIBLING of the package under test, in its repo. A repo with ecommerce/ under test gets evals/ecommerce-questions/, beside it.

Step 1: keep the file as it arrived

Copy it verbatim into evals/<set>/as-received/, and never edit that copy. Not the encoding, not the line endings, not an obvious typo.

It is there so a later disagreement is about the key and not about whether somebody transcribed the question correctly. That disagreement is guaranteed: skill:eval-loop's reference/auditing-an-answer-key.md is about what happens when it arrives, and its first check is the question against this file.

If the source cannot be committed, because it carries names, addresses, internal URLs, or row values somebody pasted into a thread, say so. Record in set.json where the original lives and a hash of it, note under sourceNote that the copy is withheld and why, and tell the user that the byte-for-byte check is now impossible for anyone without the original. Do not redact silently: a redaction you got wrong becomes the provenance record.

Step 2: classify each item by what arrived WITH the question

This is the whole job, and getting it wrong is how an unverified number gets scored as a fact.

what arrivedwritescorable on day one
the question onlya case with no goldenno. It measures coverage and discoverability, which is a real measurement
question + their query + a numbergolden.value, their query as canonicalQuery, verifiedBy: authored_query, status: provisionalafter step 3 agrees
question + a number, no querygolden.value, status: provisional, no verifiedByno. Somebody typed a number
question + prose criteriasee step 4depends on what the criterion is

Two rules over the whole table:

No golden holding a VALUE imports as verified. verified means two differently shaped derivations agreed through the truth package, and an import has performed neither. provisional goldens are unscorable by design: skill:eval-answer issues verdict: null on them. That is not a gap to work around.

It is also not a dead end, and it used to read like one. The way out is verify_goldens.py --promote, which marks a provisional golden verified once its value re-derives cleanly from the truth package and a second derivation exists. It is the only thing anywhere that writes golden.status. Tell the user that command when you hand over a set of provisionals, because until it is run the set scores nothing. A set that arrives with 60 typed numbers genuinely has 60 unverified numbers, and the number it can honestly produce on day one is a coverage score.

The exception is a golden that holds no value at all, and there are two kinds. Prose criteria (kind: criteria, step 4) have nothing to re-derive: the author's words ARE the key. So does a case whose stated answer is that the data is not there (kind: unanswerable), where the pass is a refusal that names what is missing and a confident number is the failure. Both import verified with verifiedBy: authored_criteria and score immediately. The rule is about numbers, because a number is the thing that can be wrong while looking right.

Watch for the second one in what arrives. A criterion reading "a refusal that names the missing data" is not a rubric clause on a normal case, it is the whole key, and importing it as criteria on a case with a value would make a refusal fail.

A question with no golden is a case, not a reject. Do not hold questions back until somebody derives keys. A set of bare questions already measures whether the model can express an answer at all, and the answers it produces are what the keys get derived from.

Deriving a key yourself, when nothing arrived with the question

The table above is about what ARRIVED. Sometimes nothing did, nobody knows the right number, and the set still needs keys. You may derive one, and the rules are the ones already stated rather than new ones:

  • Use your full model access and query until you can defend the number. The agent that gets tested later holds only get_context and execute_query; you are not that agent, and the gap is the point.

  • Record the query as canonicalQuery. A number with no query is somebody's guess, which is the row above it in the table.

  • Record the entities you used to get there as expectedEntities.required, in the kind:source:name form: the measures, dimensions and views the answer cannot be produced without. reference/case-format.md says to leave this field empty at import because guessing it invents a retrieval expectation nobody stated; you are not guessing, you just used them. That is the only trustworthy producer this field has, and it is what makes retrieval measurable on a set nobody sent keys for. Where the model offers two legitimate routes, write a requiredAnyOf group rather than pick one. verify_goldens.py check 5 audits every id against the model, so a wrong one is a hard finding rather than a silent retrieval miss.

  • Derive the list with the model in front of you, not from memory, and not with grep. The model knows what each field IS; an author does not reliably. Ask the question of the model: which entities does an answer to this depend on, and is each a measure or a dimension. Confirm every id with get_context before writing it. A list written from recollection is the most common way a set acquires a retrieval expectation nobody can satisfy.

    Measured: an agent given the model and one question found alternatives a hand-written key had missed on all three cases tried -- carrier name against nickname against code, destination_count against its underlying destination.airport_count. At roughly $0.13 a question. It is better than a person working from notes at finding what the model actually offers.

    It is worse at deciding what to REQUIRE, and the two rules below are why. On the same three cases it proposed demoting two named measures on the grounds that a raw expression answered identically. One of those equivalences was false; the other was true today and is not a reason to demote anything.

  • But VERIFY every "this is not strictly required" claim by running both sides. This is the rule that makes the step above safe, and it is not optional.

    The same probe got one badly wrong. Told that a named measure is not required when a raw column plus a plain aggregate answers identically, it wrote that average_plane_size was "equally answerable by avg(aircraft.aircraft_models.seats)". Executed, those are 229 and 196.93. The measure is aircraft.avg(...), an average over distinct aircraft, and the inline version averages over flights. The claim was false for exactly the measure whose entire purpose is to carry that distinction -- and it is the same mistake, on the same field, that produced the only wrong answer in the run this set came from.

    Had that proposal been written into the key, the case would have stopped testing the thing it exists to test.

    So: an equivalence is a claim about DATA, and reading a definition is not evidence. Run both expressions and compare the rows.

    But agreeing today is not a reason to demote the model's entity, and this is the part the probe got backwards. airport_count is count() returns the same 984 as a bare count() right now, and stops doing so the moment anyone adds a null filter or a deduplication to it. The key would not notice: it would go on passing while testing nothing. That applies to every measure that is currently a thin wrapper, which is most of them until the day one is not.

    requiredAnyOf is for two routes the MODEL defines -- a carrier's name and its nickname, destination_count and the destination.airport_count it wraps. Both are entities, both are retrievable, and an answer using either has gone through the semantic layer. That is a real choice and the group records it.

    It is NOT for "the model's entity, or a raw expression that happens to agree". Those are not two routes: one is the route, the other is a bypass that currently works. Keeping the named entity in required is what makes the eval measure what the semantic layer is for, and finding it IS the retrieval behaviour under test -- an agent that searches for the pre-defined entity inherits its definition and its rules, while one that rebuilds from raw columns inherits whatever it thought of.

    What the equivalence check is FOR is reading a run afterwards. When an answer came out right without the entity, comparing the two says whether it got lucky or got it wrong, and a set whose passes are merely lucky is a set of near-misses.

  • Then prove each one is retrievable, and record what finds it. check_findable.py --set <set> --mcp-url <mcp> --publisher <rest> --environment <env> --package <pkg> runs two checks and needs no model call. With --publisher it reads the COMPILED model and settles whether each field exists and what KIND it is; then it searches for each entity to confirm the index can actually deliver it.

    This matters more than it looks, because required is what BOTH retrieval numbers are computed against: whether the agent asked for the entity, and whether it came back. An id retrieval cannot deliver makes both fiction, and the case reports a retrieval miss on every run, which reads as a defect in the model or the agent rather than in the key.

    Prefer the compiled model over a grep, and prefer it in both directions. Measured on one package: dimension:flights:flight_count names a field the model really has, so grepping the text passes it, and the compiled model says flight_count is a MEASURE, which no dimension request can return. Conversely dimension:airports:own_type appears zero times in the .malloy and the compiled model declares it, because the source exposes it implicitly from the data; the grep calls it missing and acting on that deletes a good entity. verify_goldens.py check 5 is the grep, it is free and needs no server, and it is not the authority.

    A pass is a floor, not a verdict on the docs: each entity is searched by its OWN NAME, the easiest query that could find it. An entity that answers to its identifier and not to the words a question uses still fails at run time, as not retrieved.

    Writing the list forces the question an author has to answer anyway: what would a reasonable agent search for, and of what type, to find this? If you cannot state that, the agent cannot be expected to guess it.

  • status: provisional, never verified, for the reason stated above: you derived it THROUGH the model under test, so a model bug would certify its own key. verify_goldens.py --promote is still the only way out.

  • Where two readings are both defensible from the model, do not pick one. A model that documents two conventions for the same population produces two honest numbers, and choosing quietly is how a confident wrong key gets written. Go to Hold an ambiguous golden in skill:eval-loop's reference/golden-side-door.md: record the competing candidates rather than a new number.

  • What no query can settle is not a failure to derive. It is step 4's second kind, and those clauses score on day one.

Step 3: run their query, if they gave one

Do it at import, before any run. It is the cheapest finding in the whole loop.

Execute their query with execute_query against the model under test and compare with the number they stated. Three outcomes, three different things learned:

  • It returns their number. verifiedBy: authored_query. Keep status: provisional: this proves the number came from that query, and NOT that the query is right. It ran against the model under test, so a model bug certifies its own golden, which is the exact circularity the truth package exists to break. It is still far more than a typed number is worth.
  • It returns a different number. A finding on day one, before a single answer was scored. Record BOTH numbers on the case and leave the status provisional. Do not quietly adopt either one: their number may be stale, the model may have moved, or the query may have always been wrong, and which of those it is decides who owns the fix.
  • It does not compile. Also a finding, and usually the most informative one: their query names entities this model version does not have. Record the error. This is what a coverage gap looks like before anyone has phrased a question about it.

Never repair their query to make it run. A query that does not compile against this model is evidence about this model.

Show full SKILL.md (1,460 more words)Show less

Step 4: prose criteria are two different things

Decide which, per criterion. The test:

Can you write a query whose result settles it?

If yes, it is a value in disguise. "Should be about 4.2M." "The top category is Denim." "Should come back with twelve rows." There is one right answer, so the criterion is not the key, it describes one. Write it as golden.rubric, derive the value, and until it is derived the case is provisional and unscorable. Grading such a criterion as prose is how an unverified number passes: the judge reads "about 4.2M", the answer says 4.2M, and nothing ever checked whether 4.2M is right.

If no, it is a judgment about the shape of the answer, and prose is the right home for it. "Must break the total out by region." "Must say the window excludes returns." "Must not use list price." No query settles these, the judge grades them against the answer, and these cases score on day one. Write the golden as kind: criteria with no value, the criteria as golden.rubric, and verifiedBy: authored_criteria. Use golden.mustState and golden.mustNotUse where the criterion fits those fields exactly, per reference/case-format.md.

Most arriving criteria are mixed: one sentence naming a number and a shape. Split it. The number half goes provisional; the shape half scores.

Where you cannot decide, mark the case and report it rather than guessing. A criterion nobody could classify is a question for its author, and skill:eval-loop's golden side door is where it waits.

Four rules for writing the rubric itself

Each of these cost a scored case on a real set, and none of them is obvious while you are writing one.

The golden rows are the figures. The rubric's prose is a gloss on them, and where the two disagree the prose is what is stale. On one set 26 of 29 rubrics quoted a figure that appears nowhere in their own golden rows: the goldens were re-derived the next morning, the prose was not, and an agent that computed 747 and 370 -- the exact numbers in the golden JSON -- was failed against a rubric still saying 615 and 502. Say how to derive the figure, not what it equalled. verify_goldens.py check 2 reports figures in the accepting clause that are absent from the rows, but it reads only figures specific enough to be a quoted result -- a bare three-digit number is invisible to it, and 615 is exactly that. So the check is a help, not a guarantee, and the rule above is yours to keep. The judge is told the same thing from the other side: where a rubric and a golden disagree about a figure, it scores against the golden.

Do not assert the model's current behaviour. "contract_terms cannot be used here at all -- it returns ZERO rows" was true when written and false four hours later, when a commit unblocked that view. Nothing linked the two, and the judge then reasoned from the stale claim against an answer that was right. Co-locating the set with the model makes such drift visible in a diff; it does not detect it. A rubric that describes a bug is a rubric with an expiry date: write the requirement, not the defect.

When a question admits two honest populations, accept either and require the answer to name which. "Products in a category" can mean listed-in or primary-category, and on one set the two readings differed by 4,070 against 2,707. Six or more cases turned on it and one rubric had been written to accept both. That one was right. The rest were repaired afterwards, having failed correct answers in the meantime.

If the rubric accepts an alternative, the entity list must too. The two halves of a key are read by different things: the rubric is prose for the judge, expectedEntities.required is ids for retrieval scoring. They can disagree without anything noticing. A rubric saying "either the full carrier name or the nickname is fine" beside a required list naming only dimension:carriers:name scores an answer that used the nickname CORRECT and docks it recall in the same run, and that lost recall then reads as a retrieval failure. Use a requiredAnyOf group, which is satisfied when any member is delivered. verify_goldens.py reports the mismatch as a review item, but the rule is yours: it is a heuristic over prose and cannot catch every phrasing.

A trap note is not a requirement. Notes that arrive beside the questions describe what their author thought was hard, and they mention things the question never asked for. Turning one into a rubric clause invents a requirement the answerer was never given, and it happened: a term appearing nowhere in any question became a clause an answer was marked down for missing. Convert a note only where it constrains the answer to the question as asked.

Step 5: never change a question, and seal it

The question is the stimulus. It is what a human asked, and it is not yours to tidy. Do not fix a typo, normalize casing, expand an abbreviation, or sharpen a vague phrase. Vagueness is often the thing the case tests, and narrowing a question to match what an answerer keeps doing deletes the test and hands the model a pass. reference/auditing-an-answer-key.md has the audit where that happened to four questions.

At conversion, stamp questionSha: the SHA-256 of the exact question text, as written. scripts/import_cases.py --stamp does it, and refuses to overwrite a stamp that already exists.

The seal is not a derivation from the arriving file, which is why it works for an email thread as well as a CSV. It records the decision you made at conversion about what the question is. From then on, verify_goldens.py compares the stamp against the question in cases.jsonl on every audit, and a mismatch means somebody edited a question after it was imported.

A question that genuinely has to change gets a new qid, not a new stamp. A changed question is a different stimulus, and scores on the old wording must not roll into the new one.

Step 6: freeze the split

Every case gets split: dev or holdout. Diagnose and improve read dev only; the acceptance check runs both, and a set that is all dev cannot defend an accept. skill:eval-loop owns the rule; import is where it is frozen, because a split chosen after the first failures is not a holdout.

Prefer variety over volume. Cases differing in grain, source, filter shape and phrasing are what move a measurement.

Step 7: validate, and report what the set is worth

python3 scripts/import_cases.py --set evals/<set> --stamp

It checks the required fields, unique qids, a split on every case, a golden status from the allowed statuses, that nothing claims verified without the evidence for it, and that no question has drifted from its stamp. Exit 0 clean, 1 with a finding, 2 on a usage error.

Then tell the user, in these terms, what arrived:

47 cases from 50 lines
  12 scorable now
  29 provisional (21 with their query, 5 a number alone, 3 nothing to compare yet)
  6 no golden (question only)

Say what turns the 29 into scorable cases, in the same breath: a truth package (init_truth_package.py), then verify_goldens.py --set <set> --publisher <truth> --promote. A reader told only the counts has no way to know the set is not simply broken, and the provisional bucket is the one that looks like progress and is not.

The three provisional buckets are three different amounts of work, which is why they are counted apart. "With their query" needs a re-run. "A number alone" needs a derivation. "Nothing to compare yet" is a case whose criteria describe a key nobody has derived, and it is the one most easily mistaken for progress: measured on a real markdown thread, all three of its asks landed there.

Say the scorable count out loud when you report it, not just the case count. "47 cases" reads like a 47-case measurement, and 12 is the number a first run can actually score. Read the two numbers on the first line against each other too: 47 from 50 means three lines did not parse, and the findings say which.

Never

  • invent a question, or a number, or a criterion;
  • edit a question, including to fix a typo;
  • mark a golden that holds a value verified;
  • drop an item you could not classify, instead of reporting it;
  • put evals/ inside the served package tree;
  • score anything. That is skill:eval-answer.
  • skill:eval-loop: the conductor. Import is its scrape step; its reference/auditing-an-answer-key.md is the procedure for a key you doubt, and its golden side door is where an unclassifiable item waits.
  • skill:eval-answer: reference/ledger-schema.md defines every file and field this skill writes.
  • Pulling questions from production logs is the other good source, and where those logs live is a host concern. Look for a host-specific log-fetching skill; this skill takes over once you have the text.

© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/eval-import of malloydata/publisher.

  • SKILL.md
  • reference/case-format.md
  • scripts/import_cases.py
  • scripts/import_cases_test.py

Open the folder on GitHubat commit 39a546f

Compare with similar skills

Eval Import next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Import compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Import this skillmalloydata/publisher116—~5.9kAutomated safety check: PassMIT
Data Table Managern8n-io/n8n207k—~2.3kAutomated safety check: PassCustom licence
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Markitshift-labs-ai/markit1.3k—~299Automated safety check: PassMIT
Convert Fileduckdb/duckdb-skills6001 repos~720Automated safety check: NotesMIT
Research Integrity Auditxuzhougeng/wisp-science1k—~2.6kAutomated safety check: PassAGPL-3.0

Similar skills

  • Official

    Load before calling data-tables or parse-file. An agent skill from n8n-io/n8n.

    207k GitHub stars~2.3k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Markit

    shift-labs-ai/markit

    Convert files and URLs to Markdown. An agent skill from shift-labs-ai/markit.

    1.3k GitHub stars~299 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Convert File

    duckdb/duckdb-skills

    Official

    Convert any data file to another format: CSV, Parquet, JSON, Excel, GeoJSON, and more.

    600 GitHub starsUsed in 1 repo~720 tokens
    Documents & OfficeAuto-check: notes
  • Research Integrity Audit

    xuzhougeng/wisp-science

    学术审查 / research-integrity screening of a manuscript's figures and reported numbers.

    1k GitHub stars~2.6k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Jmh Benchmark Compare

    eclipse-rdf4j/rdf4j

    Parse JMH result text by finding the first header line that starts with Benchmark and contains Mode and Score, build a structured table for all columns/rows, compare overlapping benchmarks across 2+…

    420 GitHub stars~804 tokensUpdated today
    Documents & OfficeAuto-check passed

More from malloydata/publisher

All 29 skills in this repo
  • Eval Answer

    malloydata/publisher

    Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

    116 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Fix Scan Finding

    malloydata/publisher

    Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…

    116 GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Eval Loop

    malloydata/publisher

    Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

    116 GitHub stars~7.8k tokensUpdated today
    Auto-check passed
  • Eval Improve

    malloydata/publisher

    Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.

    116 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Eval Judge

    malloydata/publisher

    Decide whether ONE answer matches its golden, and say whether you believe the golden.

    116 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Malloy Analysis Report

    malloydata/publisher

    Combine validated Malloy queries into a notebook report. An agent skill from malloydata/publisher.

    116 GitHub stars~2.1k tokensUpdated today
    Auto-check passed

Questions about Eval Import

What does Eval Import do?

Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs. Eval Import is an agent skill from malloydata/publisher. Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.

When should I use Eval Import?

Eval Import fits situations like: questions arrive from outside and need to become a set; before the first run against a set nobody here authored.

How do I install Eval Import in Claude Code?

Run `npx skills add malloydata/publisher --skill eval-import -a claude-code`. Or copy the skill folder (skills/eval-import in malloydata/publisher) into .claude/skills/eval-import in your project. Claude Code loads it when a task matches its description.

How do I install Eval Import in Codex?

Run `npx skills add malloydata/publisher --skill eval-import -a codex`. Or copy the skill folder (skills/eval-import in malloydata/publisher) into .agents/skills/eval-import in your project. Codex loads it when a task matches its description.

Can I use Eval Import in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-import -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-import, .gemini/skills/eval-import, .github/skills/eval-import and .opencode/skills/eval-import in your project.

What does Eval Import need to run?

Going by SKILL.md and its folder, Eval Import needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Eval Import access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Import safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Eval Import use?

Eval Import is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Import use?

About 5.9k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Import?

Skills that share tags, products or a category with Eval Import: Data Table Manager (n8n-io/n8n, 207k stars), Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars), Markit (shift-labs-ai/markit, 1.3k stars) and Convert File (duckdb/duckdb-skills, 600 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Import?

malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 8, 2026.

Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.