Agent skill

Add Pallas Kernel

by marin-community in marin-community/marin

Add or change a named Pallas/Mosaic kernel, including its reference implementation, correctness tests, wrapper, or requested tuning.

Apache-2.0Auto-check passedTesting & QA

Install Add Pallas Kernel

skills CLI
$ npx skills add marin-community/marin --skill add-pallas-kernel -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install marin-community/marin add-pallas-kernel --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/add-pallas-kernel .claude/skills/add-pallas-kernel && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-pallas-kernel
GitHub stars
3.9k
Token cost
~1.6k tokens
SKILL.md length
721 words
Files
8
Skills in repo
41
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add or change a named Pallas/Mosaic kernel, including its reference implementation, correctness tests, wrapper, or requested tuning.

  • Works in 3 steps: Start from a reference → Write a value and gradient harness → Promote long-lived checks to pytest
  • Testing & QA work in your project
  • SKILL.md covers How to apply this skill, Kernel Deliverables, Correctness Workflow and Pallas Kernel Workflow, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Add Pallas Kernel is an agent skill from marin-community/marin. Add or change a named Pallas/Mosaic kernel, including its reference implementation, correctness tests, wrapper, or requested tuning.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files (for example `docs/api-patterns.md`, `docs/gpu-tips.md` and `docs/kernel-sources.md`).

It sits in Testing & QA. The repository describes itself as: Open-source framework for the research and development of foundation models. The licence is Apache-2.0.

When your agent uses it

  • Testing & QA work in your project

Example prompts

  • “/add-pallas-kernel”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Start from a reference
  2. Write a value and gradient harness
  3. Promote long-lived checks to pytest

What it can do on your machine

Read from SKILL.md and the folder at commit c468793. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Pallas Kernel loads about 1.6k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 721 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from marin-community/marin at commit c468793, republished under its Apache-2.0 licence (© marin-community). 721 words, ~1,560 tokens.

Download SKILL.mdSave it as .claude/skills/add-pallas-kernel/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
add-pallas-kernel
description
Add or change a named Pallas/Mosaic kernel, including its reference implementation, correctness tests, wrapper, or requested tuning.

Add or update a Pallas kernel

How to apply this skill

Load only the detail files needed for the requested work:

  • Kernel sources: read when choosing an in-repo or external kernel to imitate.
  • Performance workflow: read before benchmarking, profiling, roofline analysis, or autotuning.
  • API patterns: read before adding or changing a public kernel wrapper, fallback order, or block-size config.
  • TPU tips: read for TPU Pallas/Mosaic kernels, TPU-specific lowering failures, scoped VMEM, or TPU compiler dumps.
  • GPU tips: read for GPU Pallas/Mosaic work.
  • Deep references live under docs/reference/; read them only when the routed detail files point there.

Use research only when the user explicitly requests a multi-session research workflow.

Kernel Deliverables

For a kernel K, produce:

  • Vanilla JAX reference and Pallas wrapper with the same public API.
  • Value, gradient, CPU, and applicable accelerator parity harness.
  • Explicit backend and shape validation, with tests for ordered implementation selection and each fallback path.
  • Roofline estimate and steady-state benchmark on representative shapes/dtypes.
  • When tuning is requested, bounded autotuning, a checked-in tuned table, explicit fallback, and cached autotune-on-miss results.

Correctness Workflow

1. Start from a reference

Use an existing in-repo implementation, pseudocode, a PyTorch reference, or a JAX baseline. The baseline must be obvious and stable, not clever. If the naive baseline would materialize huge intermediates, use a streaming/blockwise baseline with identical math.

2. Write a value and gradient harness

Minimum checks:

  • Value parity over a shape/dtype grid.
  • Gradient parity on small shapes.
  • Backend numerics on CPU and accelerator backends as applicable.
  • Pointwise deviation metrics such as max/mean absolute diff, not only allclose.

Use explicit shape/dtype annotations for public APIs and references, such as jaxtyping, where available.

3. Promote long-lived checks to pytest

For in-tree kernels, add or extend tests under lib/levanter/tests/kernels/. Compare the default implementation against the reference on small CPU shapes and accelerator-aligned shapes for fast paths. Read TESTING.md and the nearest module AGENTS.md before writing or changing tests.

Pallas Kernel Workflow

Once the reference is correct, design the Pallas implementation. Use the reference as both a correctness oracle and a performance baseline.

Use existing kernels for structure and API inspiration. Read Kernel sources unless the user already named the specific kernel to follow. Unless there is a stronger local pattern, start by reimplementing the reference in Pallas.

Wrap accelerator kernel boundaries in an explicit jax.shard_map by default. This applies to pl.pallas_call, Mosaic GPU kernels, and custom FFI calls. Reshard inputs to the intended local PartitionSpec before the shard_map, keep the sequence or other nonlocal dimensions unsharded unless the kernel is explicitly written for them, and add a regression check that the lowered JAXPR or HLO contains the expected shard_map. Do not rely on XLA to infer a good sharding for an opaque kernel call boundary. Exceptions are limited to wrappers whose inputs are explicitly documented and tested as fully local or replicated.

Check correctness against the harness and reference implementation before tuning. Once the kernel is correct, run a performance harness on representative shapes/dtypes and compare against the roofline. If performance is not near the expected roofline, read Performance workflow and investigate compiler dumps, pressure signals, and tile choices before broad rewrites.

Show full SKILL.md (201 more words)Show less

API Conventions

Read API patterns before adding or changing the public wrapper, backend selection, block-size config, or input normalization contract. Keep the reference/XLA path usable even when accelerator-specific constraints are not met. Keep backend-specific validation in backend-specific modules.

Cost Estimate Requirement

Add cost_estimate= to each pl.pallas_call:

  • Use pl.estimate_cost on a body-equivalent JAX function, not a kernel body with pl.program_id.
  • Include IO bytes from call inputs/outputs.
python
from levanter.kernels.pallas.cost_estimate_utils import with_io_bytes_accessed


def _cost_estimate(
    q: jax.Array,
    k: jax.Array,
    v: jax.Array,
    *,
    kernel_inputs_specs,
    kernel_outputs_specs,
) -> pl.CostEstimate | None:
    body_cost = pl.estimate_cost(reference_impl, q, k, v)
    return with_io_bytes_accessed(
        body_cost,
        kernel_inputs_specs=kernel_inputs_specs,
        kernel_outputs_specs=kernel_outputs_specs,
    )

Definition of Done

  • Values match reference within tolerance on the tested grid.
  • Gradients match reference on small shapes.
  • CPU/reference and accelerator fast paths are covered by tests where applicable.
  • Public API, fallback semantics, block-size config, and tuned table behavior match API patterns.
  • Every Pallas, Mosaic, or FFI kernel call is inside an explicit shard_map, or its wrapper documents and tests why the inputs are fully local or replicated. Tests or profile evidence show it did not lower through unintended all-gathers.
  • Each pl.pallas_call has a reviewed cost_estimate=.
  • Benchmark/tuning artifacts include the required schema from Performance workflow.
  • Roofline performance is within expected bounds, or limitations are explicitly documented.
  • Performance improves on at least one realistic target shape, or limitations are explicitly documented.
  • Tuned table is checked in for requested hardware/shape regimes.
  • Long-running research records follow the research workflow.

© marin-community, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files in .agents/skills/add-pallas-kernel of marin-community/marin.

  • SKILL.md
  • docs/api-patterns.md
  • docs/gpu-tips.md
  • docs/kernel-sources.md
  • docs/performance-workflow.md
  • docs/reference/llo.md
  • docs/reference/profiling.md
  • docs/tpu-tips.md

Open the folder on GitHubat commit c468793

Compare with similar skills

Add Pallas Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Pallas Kernel compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Pallas Kernel this skillmarin-community/marin3.9k—~1.6kAutomated safety check: PassApache-2.0
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
Diagnosing Bugsfossasia/eventyay-interpretation1.6k32 repos~2.1kAutomated safety check: PassApache-2.0
TDDpietheinstrengholt/rssmonster56430 repos~906Automated safety check: PassMIT
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Diagnosing Bugs

    fossasia/eventyay-interpretation

    Diagnosis loop for hard bugs and performance regressions. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 32 repos~2.1k tokens
    Testing & QAAuto-check passed
  • TDD

    pietheinstrengholt/rssmonster

    Test-driven development. An agent skill from pietheinstrengholt/rssmonster.

    564 GitHub starsUsed in 30 repos~906 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • Context Driven Development

    Ibrahim-3d/orchestrator-supaconductor

    A skill your agent uses when working with Conductor's context-driven development methodology, managing project context artifacts, or understanding the relationship between product.md, tech-stack.md…

    381 GitHub starsUsed in 9 repos~2.9k tokens
    Testing & QAAuto-check passed

More from marin-community/marin

All 41 skills in this repo
  • Noslop

    marin-community/marin

    Deslop, simplify, or review low-value tests and prose only when explicitly requested for a branch or diff.

    3.9k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Use Iris

    marin-community/marin

    Use Iris to submit, inspect, debug, monitor, or recover jobs and tasks; diagnose scheduling and federation; deploy controllers; or reserve dev GPUs and TPUs.

    3.9k GitHub stars~745 tokensUpdated today
    Auto-check passed
  • Launch Rl

    marin-community/marin

    Define, validate, submit, or restart a Marin SkyRL experiment through its artifact main.

    3.9k GitHub stars~894 tokensUpdated today
    Auto-check passed
  • Marina Applet

    marin-community/marin

    Build, validate, publish, update, inspect, query, roll back, or archive a dynamic Marina applet.

    3.9k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Query Finelog

    marin-community/marin

    Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding.

    3.9k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Trace Pulumi Diff

    marin-community/marin

    Run a read-only preview for a specified Marin infra/pulumi stack and trace each pending resource change to merged pull requests since its latest successful update when that update records a clean…

    3.9k GitHub stars~663 tokensUpdated today
    Auto-check passed

Categories

Questions about Add Pallas Kernel

What does Add Pallas Kernel do?

Add or change a named Pallas/Mosaic kernel, including its reference implementation, correctness tests, wrapper, or requested tuning. Add Pallas Kernel is an agent skill from marin-community/marin. Add or change a named Pallas/Mosaic kernel, including its reference implementation, correctness tests, wrapper, or requested tuning.

When should I use Add Pallas Kernel?

Add Pallas Kernel fits situations like: testing & QA work in your project.

How do I install Add Pallas Kernel in Claude Code?

Run `npx skills add marin-community/marin --skill add-pallas-kernel -a claude-code`. Or copy the skill folder (.agents/skills/add-pallas-kernel in marin-community/marin) into .claude/skills/add-pallas-kernel in your project. Claude Code loads it when a task matches its description.

How do I install Add Pallas Kernel in Codex?

Run `npx skills add marin-community/marin --skill add-pallas-kernel -a codex`. Or copy the skill folder (.agents/skills/add-pallas-kernel in marin-community/marin) into .agents/skills/add-pallas-kernel in your project. Codex loads it when a task matches its description.

Can I use Add Pallas Kernel in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add marin-community/marin --skill add-pallas-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-pallas-kernel, .gemini/skills/add-pallas-kernel, .github/skills/add-pallas-kernel and .opencode/skills/add-pallas-kernel in your project.

What does Add Pallas Kernel need to run?

SKILL.md names no scripts, command-line tools or credentials: Add Pallas Kernel is instructions for the agent only. Our summary lists: Python 3.

Does Add Pallas Kernel access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Add Pallas Kernel safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Pallas Kernel use?

Add Pallas Kernel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Pallas Kernel use?

About 1.6k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Pallas Kernel?

Skills that share tags, products or a category with Add Pallas Kernel: Web Application Testing (anthropics/skills, 180k stars), Diagnosing Bugs (fossasia/eventyay-interpretation, 1.6k stars), TDD (pietheinstrengholt/rssmonster, 564 stars) and TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Pallas Kernel?

marin-community (a GitHub organization) maintains it in marin-community/marin, which has 3,921 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 10, 2026.

Source: marin-community/marin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.