Agent skill

Colm Artifact Evaluation

by brycewang-stanford in brycewang-stanford/Awesome-Journal-Skills

A skill your agent uses when packaging the artifacts of a COLM paper — model weights, training data, prompts, evaluation sets, and cached model outputs — for anonymous review and public…

MITAuto-check passedLegal & Compliance

Install Colm Artifact Evaluation

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill colm-artifact-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills colm-artifact-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/COLM-Skills/skills/colm-artifact-evaluation .claude/skills/colm-artifact-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
colm-artifact-evaluation
GitHub stars
1.2k
Token cost
~1.8k tokens
SKILL.md length
871 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when packaging the artifacts of a COLM paper — model weights, training data, prompts, evaluation sets, and cached model outputs — for anonymous review and public…

  • Works in 5 steps: De-anonymize the repository; restore… → Upload weights/data with cards and… → Add the paper's OpenReview URL to every… → …
  • Packaging the artifacts of a COLM paper — model weights
  • SKILL.md covers The COLM artifact taxonomy, One-command reproduction, Anonymous review packaging and Datasheets for anything you…, plus 4 more sections
  • Calls make

What it does

Colm Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging the artifacts of a COLM paper — model weights, training data, prompts, evaluation sets, and cached model outputs — for anonymous review and public post-acceptance release, navigating licenses, API terms-of-service limits, and the absence of a formal COLM artifact track.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Legal & Compliance, covering Policy and terms drafting. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Packaging the artifacts of a COLM paper — model weights
  • Evaluation sets
  • Cached model outputs — for anonymous review and public post-acceptance release
  • Navigating licenses

Example prompts

  • “/colm-artifact-evaluation”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. De-anonymize the repository; restore real hub org names and W&B links.
  2. Upload weights/data with cards and licenses; tag the exact code release cited in
  3. Add the paper's OpenReview URL to every artifact so provenance points both ways.
  4. Freeze a colm2026 git tag — the paper cites a snapshot, not a moving main.
  5. Update the paper's availability statement to match what actually shipped

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Colm Artifact Evaluation loads about 1.8k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 871 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 871 words, ~1,814 tokens.

Download SKILL.mdSave it as .claude/skills/colm-artifact-evaluation/SKILL.md (or your agent's skills folder).
name
colm-artifact-evaluation
description
Use when packaging the artifacts of a COLM paper — model weights, training data, prompts, evaluation sets, and cached model outputs — for anonymous review and public post-acceptance release, navigating licenses, API terms-of-service limits, and the absence of a formal COLM artifact track.

COLM Artifact Evaluation

COLM had no formal artifact-evaluation track verifiable for the 2026 cycle (checked 2026-07-08; 待核实 each edition). That absence does not lower the bar — it moves the audit into ordinary review, where artifact quality influences scores without a rubric to appeal to. Package as if a skeptical reviewer will spend ten minutes with your materials, because at this venue one usually will.

The COLM artifact taxonomy

LM papers produce artifact types with very different release mechanics; inventory yours before deciding anything:

ArtifactReview-time formRelease-time formBlocking question
Code (training/eval)Anonymized repo or supplement ZIPPublic repo, tagged releaseDoes one command reproduce one table?
PromptsVerbatim appendix + files in packageSame, publicExact strings, incl. system prompts?
Fine-tuned weightsUsually described, not uploaded (size)Model hub upload with model cardDoes the base model's license permit derivative release?
Training/eval data you builtAnonymized sample + datasheetFull release with licenseAny personal data, scraped ToS conflicts, or annotator-privacy issues?
Cached model outputsSample in supplementFull archiveDoes the provider's ToS permit publishing outputs at this scale?
Human-eval materialsInstructions + interface in appendixSameIRB/consent status stated?

The two questions authors most often skip are in the right-hand column: derivative weight licensing and output-publication ToS. Both can void a promised release after acceptance — resolve them before the paper commits to anything.

One-command reproduction

The credibility core of the package is a single entry point per headline result:

makefile
# Makefile at package root — one target per main-text table/figure
table2:            ## headline comparison, ~40 GPU-min or ~$8 API spend
	python run_eval.py --config eval/run-042.yaml --out results/table2.csv
figure3:           ## scaling curve from cached outputs (no model access needed)
	python plots/scaling.py --cache outputs/cache.jsonl --out figs/figure3.pdf
verify:            ## regenerate all numbers from cached outputs only
	python verify_from_cache.py --tolerance 0.1

The verify-from-cache target is the LM-specific trick: reviewers without GPUs or API budgets can still confirm that your published numbers follow from your recorded model responses. It converts "trust me" into "check me" at zero compute cost, and it keeps working after API models drift or deprecate (colm-reproducibility).

Anonymous review packaging

  • Build from a clean export and grep for identity leaks — the commands and channel list live in colm-supplementary; org-scoped model-hub IDs are the leak class unique to LM work.
  • If weights must be inspectable at review time, an anonymized hub account or a size-reduced distilled checkpoint are the workable options; a link to your lab's account is a double-blind violation under COLM's no-identifying-links rule.
  • Include a MANIFEST.md: inventory, license per item, compute needed per target, and what is deliberately absent with the reason ("training corpus omitted: contains licensed text; filtering scripts included instead").

Datasheets for anything you release

For each dataset or evaluation set: how items were created (author population, LLM-generated fraction — disclosable under the 2026 LLM policy), collection dates (this doubles as contamination documentation for future users), license and source licenses, known biases and coverage gaps, and a contamination canary if you want future training runs to be detectable. For model releases: intended use, evaluation scope, and known failure modes. These documents are cheap at packaging time and impossible to reconstruct honestly later.

Post-acceptance release sequence

  1. De-anonymize the repository; restore real hub org names and W&B links.
  2. Upload weights/data with cards and licenses; tag the exact code release cited in the camera-ready.
  3. Add the paper's OpenReview URL to every artifact so provenance points both ways.
  4. Freeze a colm2026 git tag — the paper cites a snapshot, not a moving main.
  5. Update the paper's availability statement to match what actually shipped (colm-camera-ready owns the deadline: August 7, 2026 this cycle).
Show full SKILL.md (327 more words)Show less

The ten-minute reviewer walkthrough

Dry-run the package as the busiest plausible reviewer before every upload. The sequence they follow is predictable, so optimize for it in order:

  1. Minute 1-2: open MANIFEST.md. If there is no manifest, they grep for a README and form their opinion from whatever half-stale file they find. The manifest is the cheapest score you will ever buy.
  2. Minute 3-4: look for the entry point. make table2 or an equivalent single command. If the first thing they see is a 400-line setup guide with cluster assumptions, the walkthrough ends here.
  3. Minute 5-7: run verify from cache. No GPUs, no keys, under a minute of compute — this is the step that actually gets executed in practice, which is why the cached-outputs archive earns its place in the package.
  4. Minute 8-9: spot-check a prompt file against the paper's appendix. Any mismatch between packaged prompts and printed prompts contaminates trust in both.
  5. Minute 10: skim the code for identity leaks and hardcoded secrets — partly ethics-duty, partly curiosity. This is where an org-scoped from_pretrained string ends your anonymity.

Sizing note: keep the review package lean — cached outputs can be sampled down to what verify needs, with the full archive promised for release. No supplementary size cap was verifiable for COLM 2026 (待核实), but a multi-gigabyte upload fails socially even where it succeeds technically.

Longevity: the two-year test

A COLM artifact's real audience arrives later — the group in two years trying to compare against you. Two cheap investments serve them: freeze an environment manifest (exact package versions; a container digest if you can), and write the MANIFEST.md assuming every external URL in it will eventually rot — name artifacts by content hash where possible so mirrors stay verifiable. The venue publishes on OpenReview, where your artifact links are permanently attached to the paper's public page; links that die quietly are the failure mode, so prefer archival hosts for anything you would want cited.

Output format

text
[Inventory] code ▢ prompts ▢ weights ▢ data ▢ cached-outputs ▢ human-eval ▢
[Legal gates] base-model license: <ok/blocks release>  output-ToS: <ok/limits>
[One-command check] targets exist for: <tables/figures>  verify-from-cache: ▢
[Anonymity] package clean / leaks: <items>
[Manifest + datasheets] present / missing: <which>
[Release plan] <ordered steps with dates>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in COLM-Skills/skills/colm-artifact-evaluation of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Colm Artifact Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Colm Artifact Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Colm Artifact Evaluation this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.8kAutomated safety check: PassMIT
Master Agreement Generatoraffaan-m/ECC276k—~2.9kAutomated safety check: PassMIT
Pii Contract Analyzegregmos/PII-Shield149—~8.9kAutomated safety check: NotesMIT
Privacy Eukimlawtech/korean-privacy-terms586—~968Automated safety check: PassApache-2.0
Agents In The Teamborghei/Claude-Skills886—~4.2kAutomated safety check: PassMIT
Terms Of Service Generatorzubair-trabzada/ai-legal-claude1.8k—~2.9kAutomated safety check: PassNone

Similar skills

  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    276k GitHub stars~2.9k tokensUpdated 4 days ago
    Legal & ComplianceAuto-check passed
  • Pii Contract Analyze

    gregmos/PII-Shield

    Universal legal document processor with PII anonymization. An agent skill from gregmos/PII-Shield.

    149 GitHub stars~8.9k tokensUpdated 3 mo ago
    Legal & ComplianceAuto-check: notes
  • Privacy Eu

    kimlawtech/korean-privacy-terms

    EU 사용자 대상 서비스용 Privacy Notice·Terms of Service·Consent Modal·Cookie Banner 자동 생성.

    586 GitHub stars~968 tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    886 GitHub stars~4.2k tokensUpdated 2 days ago
    Legal & ComplianceAuto-check passed
  • Terms Of Service Generator

    zubair-trabzada/ai-legal-claude

    Generates complete, GDPR/CCPA-compliant Terms of Service for a website or SaaS product, with plain English summaries for each section

    1.8k GitHub stars~2.9k tokensUpdated 6 mo ago
    Legal & ComplianceAuto-check passed
  • Tos Clause Scanner

    zebbern/claude-code-guide

    Audit Terms of Service, user agreements, and privacy policies for consumer risks, producing a structured report that flags unfair clauses, data traps, and liability issues.

    4.7k GitHub starsUsed in 1 repo~3.3k tokens
    Legal & ComplianceAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 12 days ago
    Auto-check passed

Questions about Colm Artifact Evaluation

What does Colm Artifact Evaluation do?

A skill your agent uses when packaging the artifacts of a COLM paper — model weights, training data, prompts, evaluation sets, and cached model outputs — for anonymous review and public…. Colm Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging the artifacts of a COLM paper — model weights, training data, prompts, evaluation sets, and cached model outputs — for anonymous review and public post-acceptance release, navigating licenses, API terms-of-service limits, and the absence of a formal COLM artifact track.

When should I use Colm Artifact Evaluation?

Colm Artifact Evaluation fits situations like: packaging the artifacts of a COLM paper — model weights; evaluation sets; cached model outputs — for anonymous review and public post-acceptance release; navigating licenses.

How do I install Colm Artifact Evaluation in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill colm-artifact-evaluation -a claude-code`. Or copy the skill folder (COLM-Skills/skills/colm-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/colm-artifact-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Colm Artifact Evaluation in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill colm-artifact-evaluation -a codex`. Or copy the skill folder (COLM-Skills/skills/colm-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/colm-artifact-evaluation in your project. Codex loads it when a task matches its description.

Can I use Colm Artifact Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill colm-artifact-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/colm-artifact-evaluation, .gemini/skills/colm-artifact-evaluation, .github/skills/colm-artifact-evaluation and .opencode/skills/colm-artifact-evaluation in your project.

What does Colm Artifact Evaluation need to run?

Going by SKILL.md and its folder, Colm Artifact Evaluation needs the command-line tools its instructions call (make). Our summary lists: Python 3.

Does Colm Artifact Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Colm Artifact Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Colm Artifact Evaluation use?

Colm Artifact Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Colm Artifact Evaluation use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Colm Artifact Evaluation?

Skills that share tags, products or a category with Colm Artifact Evaluation: Master Agreement Generator (affaan-m/ECC, 276k stars), Pii Contract Analyze (gregmos/PII-Shield, 149 stars), Privacy Eu (kimlawtech/korean-privacy-terms, 586 stars) and Agents In The Team (borghei/Claude-Skills, 886 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Colm Artifact Evaluation?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,228 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.