Agent skill

Malloy Document

by malloydata in malloydata/publisher

Add documentation with (doc) tags to Malloy models so fields and sources are described in plain language.

MITAuto-check passedWriting & Content

Install Malloy Document

skills CLI
$ npx skills add malloydata/publisher --skill malloy-document -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install malloydata/publisher malloy-document --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/malloy-document .claude/skills/malloy-document && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
malloy-document
GitHub stars
116
Token cost
~2.7k tokens
SKILL.md length
1,134 words
Files
1
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

Add documentation with (doc) tags to Malloy models so fields and sources are described in plain language.

  • User asks to add documentation
  • SKILL.md covers #(doc) Tag, #(filter): deprecated, see…, internal: and private::… and Annotating Columns in Include…, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Document the model

What it does

Malloy Document is an agent skill from malloydata/publisher. Add documentation with (doc) tags to Malloy models so fields and sources are described in plain language. Use when user asks to "add documentation", "add doc tags", "document the model", or wants fields and sources described for natural-language search and discovery. For runtime filters, declare given: parameters (see the malloy-model skill); (filter) is deprecated. Filters are a runtime/modeling construct (governance, latency, correctness), not a documentation tag.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Writing & Content, covering Plain language and style rules. The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.

When your agent uses it

  • User asks to add documentation
  • Document the model
  • Wants fields and sources described for natural-language search and discovery

Example prompts

  • “add documentation”
  • “add doc tags”
  • “document the model”
  • “/malloy-document”

What it can do on your machine

Read from SKILL.md and the folder at commit acc1acd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are malloy).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Malloy Document loads about 2.7k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 1,134 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from malloydata/publisher at commit acc1acd, republished under its MIT licence (© malloydata). 1,134 words, ~2,655 tokens.

Download SKILL.mdSave it as .claude/skills/malloy-document/SKILL.md (or your agent's skills folder).
name
malloy-document
description
Add documentation with #(doc) tags to Malloy models so fields and sources are described in plain language. Use when user asks to "add documentation", "add doc tags", "document the model", or wants fields and sources described for natural-language search and discovery. For runtime filters, declare given: parameters (see the malloy-model skill); #(filter) is deprecated. Filters are a runtime/modeling construct (governance, latency, correctness), not a documentation tag.
<!--
Copyright (c) Credible Data Inc.
SPDX-License-Identifier: MIT
-->

Documenting a Malloy Model

Add #(doc) tags to describe sources and fields in plain language so they are easy to find and understand:

TagPurposeGoes on
#(doc)Plain-language description for natural-language searchsource, dimension, measure, view, join
#(filter)Deprecated. Never add one; declare a given: instead (see malloy-model)source

#(doc) is a standard Malloy annotation. It documents a field or source with a human-readable description that downstream tools can surface and search against.

#(doc) Tag

Add before any source, dimension, measure, view, or join. When multiple fields share a keyword, use it once as a block header. Tags and field names are indented under the keyword; tags go on the line(s) directly above the field they annotate.

Tag ordering (when a field has multiple tags): #(doc) → render tags (# currency, # label, etc.) → field name. Separate each field group with a blank line:

malloy
#(doc) Customer who placed the order
join_one: users with user_id

dimension:
  #(doc) Date the order was placed (UTC)
  order_date is created_at::date

measure:
  #(doc) Total revenue from all orders in USD
  # currency
  revenue is sum(total)
Writing Doc Strings for Retrieval

Doc strings power natural-language search: users type plain-English questions and the system matches against your #(doc) strings. Write descriptions that match how analysts would search:

  • Include business meaning, not code mechanics: what it represents, not how it's implemented
  • Include units (USD, count, percentage): a unit is part of what a number means. For counts, name the unit being counted and say whether it counts distinct entities or events: "total students enrolled" on a subject×term-grain measure counts enrolments, not students, and a student taking four subjects counts four times. If the model cannot answer the distinct-entity version, say so in the doc.
  • List a categorical field's values only while the list stays short (roughly ten or fewer). A handful of values makes a description concrete; past that, say what the field captures instead, because the dump crowds out the meaning and goes stale the moment someone adds a value. Treat ten as a rule of thumb, not a hard cap.
  • Avoid Malloy jargon: never use "filterable", "groupable", "dimension", "measure", "aggregation"

Good examples:

  • #(doc) Total revenue from completed orders in USD matches "what was our revenue?"
  • #(doc) Customer signup date (UTC) matches "when did the customer join?"
  • #(doc) Order status: pending, processing, shipped, delivered, cancelled matches "what are the order statuses?"

Bad examples:

  • #(doc) Filterable dimension for order status: no analyst searches for "filterable"
  • #(doc) Groupable by region: "groupable" is a system concept
  • #(doc) Aggregation of total sales: "aggregation" doesn't match natural queries
Mark conventions as conventions

A #(doc) must let a reader tell a measured fact from a choice someone made. Any dimension or measure encoding a threshold, bucket boundary, or business definition that the user did not explicitly confirm must say so in its own doc string:

malloy
// WRONG - a chosen cutoff stated as fact
#(doc) Popularity band: Hit (70+), Popular (40-69), Moderate (15-39), Obscure (<15)

// RIGHT - the choice is visible and auditable
#(doc) Popularity band: Hit (70+), Popular (40-69), Moderate (15-39), Obscure (<15).
#(doc) 70 follows the source dataset's own high-popularity cutoff; 40 and 15 are
#(doc) working boundaries for this model, not settled by the data.

These are governed models: a threshold nobody confirmed is an assumption, and an unlabeled assumption reads as a fact to everyone downstream, including the agents that answer questions from these docs. The hedge in the #(doc) is the artifact-time record; the "Flag Ambiguous Descriptions" table below is the conversation-time surface for getting them confirmed, and modeling-notes.md's "Open decisions" section (see skill:malloy-modeling) is where they wait for a subject-matter expert.

Do not hedge measured facts: avg_energy is avg(energy) needs no caveat. Hedge only where a domain expert could reasonably choose differently.

#(filter): deprecated, see malloy-model

#(filter) is deprecated. Never add a #(filter) annotation: every use, including required, implicit, and date/number ranges, has a given: form. See skill:malloy-model § Parameterizing sources with given:.

#(filter) is also a #(...)-shaped annotation, but unlike #(doc) it's a runtime/modeling construct: it shapes governance, query latency, and correctness, not discoverability. How to read and migrate an existing one lives in malloy-model § Legacy: reading an existing #(filter) model.

One rule worth knowing here: a filter the model reads lives on the source, never on the consumer. Ad-hoc reports and notebooks that import a source inherit its givens automatically, and import them by name rather than re-declaring them. A notebook or dashboard may declare a given: of its own for a control only its own tiles read, and nothing else.

Show full SKILL.md (480 more words)Show less

internal: and private:: column-level access in a source

#(doc) describes what's exposed. Two access modifiers control what's exposed in the first place, and both live inside a source's include {} block. They are about the source's public API and data sensitivity, not about documentation, so reach for them when curating which columns callers can pick.

MechanismLayerWhy you reach for it
internal:Inside a source (one column in include {})The column isn't part of your model's public API. Common reasons: data is messy (empty/garbage, raw JSON, duplicates), or a documented derived dimension already supersedes it, or the raw column exists only to be joined on / referenced internally and shouldn't appear as a dimension callers can pick. The data may be perfectly fine, it's just not what you want exposed.
private:Inside a source (one column in include {})The data is sensitive: SSN, raw credit card, password. Governance / security concern; a harder block than internal:.

In one sentence: internal: and private: shape what's inside a source's public API; #(doc) describes the fields you do expose.

Example

A base source pulled from a messy raw table often uses internal: to drop raw fields from the public API, while documenting the curated columns with #(doc).

malloy
// orders_base.malloy
#(doc) Raw orders. Use orders.malloy as the entry point for analysis.
source: orders_base is conn.table('orders_raw')
  include {
    public: id, customer_id, order_date, total
    internal: raw_json_payload, deprecated_status_code, _temp_dedup_marker
  }
  extend {
    primary_key: id
  }
malloy
// orders.malloy
import "orders_base.malloy"

#(doc) Order analysis. Use for revenue, fulfillment, and customer-order joins.
source: orders is orders_base extend {
  // joins, measures, curated dimensions
}

The base source stays fully queryable (run: orders_base -> { ... } still works); internal: only governs which columns appear as public dimensions callers can pick.

Annotating Columns in Include (Experimental)

With ##! experimental.access_modifiers, you can add #(doc) tags to raw table columns inside include blocks. This documents columns without redefining them as dimensions.

malloy
##! experimental.access_modifiers

source: orders is conn.table('orders') include {
  public:
    #(doc) Order line item identifier
    id

    #(doc) Customer email address
    email

    #(doc) Order status: pending, shipped, delivered
    status

  // internal: only for verified noise (empty cols, raw JSON blobs, duplicates)
}
extend {
  // ... dimensions and measures
}

When to use:

  • Documenting raw columns without creating explicit dimensions
  • Curating which columns are public vs internal

Source-Level Documentation

Document when to use a source, not what it contains. Dimensions and measures can already be searched directly, so the source-level #(doc) should describe what questions/analyses this source answers.

Base source files: Document what the table represents.

malloy
#(doc) Customer records with demographics and segmentation. One row per customer.
source: customers is conn.table('sales.customers') extend { ... }

Source files: Document what analytical questions the source answers.

malloy
#(doc) Customer health analysis. Use for retention, segmentation, churn risk, and lifetime value. For order-level analysis, use order_analysis instead.
source: customer_health is customers extend { ... }

Best practices:

  • Add #(doc) to all base source and joined source definitions
  • Base source docs: describe what the table is (one row per what)
  • Source docs: describe what questions/analyses the source answers
  • Documentation happens per-source-file, not in one monolithic file

Flag Ambiguous Descriptions

After writing #(doc) tags, present any that required judgment to the user for confirmation:

FieldProposed docConfidenceUncertainty
total"Total order amount in USD"MediumCould be gross or net, verified with sample query
status"Order status: pending, shipped, delivered"HighValues confirmed via a query of distinct values

Only flag fields where the description required assumptions about business meaning, units, or valid values. When in doubt about valid values, run a quick query against the data to confirm them before writing the description. Use get_context to ground yourself in the package's sources and fields and execute_query to check distinct values, for example run: source -> { group_by: status }.

Done

Step complete. Output: #(doc) tags added to all public fields and sources.

© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/malloy-document of malloydata/publisher.

Open the folder on GitHubat commit acc1acd

Compare with similar skills

Malloy Document next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Malloy Document compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Malloy Document this skillmalloydata/publisher116—~2.7kAutomated safety check: PassMIT
Asd Ste100danyuchn/asd-ste100-skill3.9k—~4.1kAutomated safety check: PassMIT
Natural Japanese Business Writingcoji/natural-japanese1.9k—~2.1kAutomated safety check: PassMIT
Technical Writing Standardcursor/plugins10k10 repos~2.4kAutomated safety check: PassNone
PgjevrealZachi/pg-jev1k—~2.9kAutomated safety check: PassCustom licence
Defensive Writing Editorlennney/stop-that-shit2.5k1 repos~902Automated safety check: PassMIT

Similar skills

  • Asd Ste100

    danyuchn/asd-ste100-skill

    A skill your agent uses when English text must be parsed without a human to resolve ambiguity — tool descriptions, error messages, inter-agent instructions, system prompts, status reports — and…

    3.9k GitHub stars~4.1k tokensUpdated 3 days ago
    Writing & ContentAuto-check passed
  • Writes and edits Japanese business documents so they read clearly and naturally, removes AI-sounding phrasing and can score how AI-like a text reads.

    1.9k GitHub stars~2.1k tokensUpdated 1 mo ago
    Writing & ContentAuto-check passed
  • Official

    Applies four layers of technical-writing rules to docs, RFCs, readmes, PR descriptions and commit messages so a tired engineer follows them on the first read.

    10k GitHub starsUsed in 10 repos~2.4k tokens
    Writing & ContentAuto-check passed
  • Pgjev

    realZachi/pg-jev

    Install, configure, query and explain pgjev (the jev PostgreSQL extension that filters, ranks and classifies rows with plain-language conditions via TypeSafe's Jev model).

    1k GitHub stars~2.9k tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Defensive Writing Editor

    lennney/stop-that-shit

    Cuts defensive disclaimers, stacked hedging and self-protective narration from proposals and summaries, keeping only limits that affect the reader's decision.

    2.5k GitHub starsUsed in 1 repo~902 tokens
    Writing & ContentAuto-check passed
  • C Code Style

    MaJerle/c-code-style

    Write, edit, or review C (and C-compatible header) code according to C coding style rules, then check and run clang-format to enforce formatting.

    1.4k GitHub stars~1.3k tokensUpdated 28 days ago
    Writing & ContentAuto-check passed

More from malloydata/publisher

All 29 skills in this repo
  • Eval Answer

    malloydata/publisher

    Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

    116 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Fix Scan Finding

    malloydata/publisher

    Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…

    116 GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Eval Import

    malloydata/publisher

    Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.

    116 GitHub stars~5.9k tokensUpdated today
    Auto-check passed
  • Eval Loop

    malloydata/publisher

    Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

    116 GitHub stars~7.8k tokensUpdated today
    Auto-check passed
  • Eval Improve

    malloydata/publisher

    Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.

    116 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Eval Judge

    malloydata/publisher

    Decide whether ONE answer matches its golden, and say whether you believe the golden.

    116 GitHub stars~3.4k tokensUpdated today
    Auto-check passed

Questions about Malloy Document

What does Malloy Document do?

Add documentation with (doc) tags to Malloy models so fields and sources are described in plain language. Malloy Document is an agent skill from malloydata/publisher. Add documentation with (doc) tags to Malloy models so fields and sources are described in plain language.

When should I use Malloy Document?

Malloy Document fits situations like: user asks to add documentation; document the model; wants fields and sources described for natural-language search and discovery.

How do I install Malloy Document in Claude Code?

Run `npx skills add malloydata/publisher --skill malloy-document -a claude-code`. Or copy the skill folder (skills/malloy-document in malloydata/publisher) into .claude/skills/malloy-document in your project. Claude Code loads it when a task matches its description.

How do I install Malloy Document in Codex?

Run `npx skills add malloydata/publisher --skill malloy-document -a codex`. Or copy the skill folder (skills/malloy-document in malloydata/publisher) into .agents/skills/malloy-document in your project. Codex loads it when a task matches its description.

Can I use Malloy Document in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill malloy-document -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/malloy-document, .gemini/skills/malloy-document, .github/skills/malloy-document and .opencode/skills/malloy-document in your project.

What does Malloy Document need to run?

SKILL.md names no scripts, command-line tools or credentials: Malloy Document is instructions for the agent only.

Does Malloy Document access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Malloy Document safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Malloy Document use?

Malloy Document is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Malloy Document use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Malloy Document?

Skills that share tags, products or a category with Malloy Document: Asd Ste100 (danyuchn/asd-ste100-skill, 3.9k stars), Natural Japanese Business Writing (coji/natural-japanese, 1.9k stars), Technical Writing Standard (cursor/plugins, 10k stars) and Pgjev (realZachi/pg-jev, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Malloy Document?

malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 7, 2026.

Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.