Agent skill

Areno Develop Kernel

by inclusionAI in inclusionAI/AReno

Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Areno Develop Kernel

skills CLI
$ npx skills add inclusionAI/AReno --skill areno-develop-kernel -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install inclusionAI/AReno areno-develop-kernel --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/inclusionAI/AReno.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/areno-develop-kernel .claude/skills/areno-develop-kernel && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
areno-develop-kernel
GitHub stars
323
Token cost
~498 tokens
SKILL.md length
173 words
Files
5 (incl. scripts, references)
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator.

  • Works in 7 steps: Locate Python wrapper, extension… → Add a small PyTorch reference and… → Implement forward and backward before… → …
  • Changes touch areno/accel
  • SKILL.md covers Generic harnesses and Workflow
  • Runs Python scripts from its folder; calls python and pip

What it does

Areno Develop Kernel is an agent skill from inclusionAI/AReno. Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator. Use when changes touch areno/accel or a runtime kernel boundary and require forward, backward, dtype, layout, CUDA graph, or benchmark validation.

Its SKILL.md is about 500 tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/kernel-checklist.md` and `scripts/benchmark_operator.py`).

It sits in AI & LLM Engineering. It works with CUDA. The repository describes itself as: An easy-to-use, fast toolkit to scale up RL post-training on a single node. The licence is Apache-2.0.

When your agent uses it

  • Changes touch areno/accel
  • A runtime kernel boundary and require forward
  • Benchmark validation

Example prompts

  • “/areno-develop-kernel”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Locate Python wrapper, extension registration, C++/CUDA source, engine layer, and model call site.
  2. Add a small PyTorch reference and shape/dtype assertions.
  3. Implement forward and backward before using it in training.
  4. Test representative shapes, boundary shapes, supported non-contiguous layouts, finite output, and gradients.
  5. Validate TP/sequence-parallel local shapes and CUDA graph capture/replay.
  6. After pulling a branch that changes areno/accel, rebuild remotely with pip install -e . --no-deps --no-build-isolation. Do not reinstall…
  7. Benchmark only after correctness. Read references/kernel-checklist.md.

What it can do on your machine

Read from SKILL.md and the folder at commit ce35e0d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Areno Develop Kernel loads about 498 tokens when it runs, and up to ~664 if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 173 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~498
With references · SKILL.md plus every file in references/, read only if the agent opens them
~664

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from inclusionAI/AReno at commit ce35e0d, republished under its Apache-2.0 licence (© inclusionAI). 173 words, ~498 tokens.

Download SKILL.mdSave it as .claude/skills/areno-develop-kernel/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
areno-develop-kernel
description
Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator. Use when changes touch areno/accel or a runtime kernel boundary and require forward, backward, dtype, layout, CUDA graph, or benchmark validation.

Develop an AReno Kernel

Define the mathematical reference, shapes, dtype, layout, supported devices, and backward contract before implementation.

Develop and commit on a local branch. A remote GPU checkout is validation-only: pull the committed branch there, and never edit or copy source files into it.

Generic harnesses

Both callables must accept one tensor and return one tensor:

bash
python .agents/skills/areno-develop-kernel/scripts/check_operator.py \
  --reference package.module:reference --candidate package.module:candidate \
  --shape 8,16 --dtype float32 --device cuda

python .agents/skills/areno-develop-kernel/scripts/benchmark_operator.py \
  --callable package.module:candidate --shape 8,16 --dtype float16 --device cuda

For multi-input or stateful kernels, extend a focused test under tests/ rather than weakening this harness.

Workflow

  1. Locate Python wrapper, extension registration, C++/CUDA source, engine layer, and model call site.
  2. Add a small PyTorch reference and shape/dtype assertions.
  3. Implement forward and backward before using it in training.
  4. Test representative shapes, boundary shapes, supported non-contiguous layouts, finite output, and gradients.
  5. Validate TP/sequence-parallel local shapes and CUDA graph capture/replay.
  6. After pulling a branch that changes areno/accel, rebuild remotely with pip install -e . --no-deps --no-build-isolation. Do not reinstall for Python-only changes.
  7. Benchmark only after correctness. Read references/kernel-checklist.md.

Do not add a silent production fallback for a required kernel. Report unsupported cases explicitly.

© inclusionAI, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in .agents/skills/areno-develop-kernel of inclusionAI/AReno.

  • SKILL.md
  • agents/openai.yaml
  • references/kernel-checklist.md
  • scripts/benchmark_operator.py
  • scripts/check_operator.py

Open the folder on GitHubat commit ce35e0d

Compare with similar skills

Areno Develop Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Areno Develop Kernel compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Areno Develop Kernel this skillinclusionAI/AReno323—~498Automated safety check: PassApache-2.0
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Benchmark TuneMesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0
Cuda Kernel OptimizerKernelFlow-ops/cuda-optimized-skill212—~4.3kAutomated safety check: PassMIT
DGX Spark Training Gotchaswshobson/agents40k1 repos~2kAutomated safety check: PassMIT

Similar skills

  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Benchmark Tune

    Mesh-LLM/mesh-llm

    A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

    3.5k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cuda Kernel Optimizer

    KernelFlow-ops/cuda-optimized-skill

    Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

    212 GitHub stars~4.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed

More from inclusionAI/AReno

All 10 skills in this repo
  • Areno Debug Runtime

    inclusionAI/AReno

    Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno.

    323 GitHub stars~486 tokensUpdated 13 days ago
    Auto-check passed
  • Areno Run Training

    inclusionAI/AReno

    Run, configure, retry, and validate AReno SFT, DPO, GSPO, GRPO, PPO, and agentic training.

    323 GitHub stars~782 tokensUpdated 13 days ago
    Auto-check passed
  • Areno Add Algorithm

    inclusionAI/AReno

    Add or modify an AReno algorithm, trainer, loss, advantage calculation, role model, or algorithm-specific configuration.

    323 GitHub stars~501 tokensUpdated 13 days ago
    Auto-check passed
  • Areno Profile Performance

    inclusionAI/AReno

    Measure and diagnose AReno rollout, prefill, decode, training, checkpoint, role-switch, communication, or Python scheduling performance.

    323 GitHub stars~848 tokensUpdated 13 days ago
    Auto-check passed
  • Compare an AReno branch, model, checkpoint, algorithm, scheduler, or kernel against a baseline.

    323 GitHub stars~389 tokensUpdated 13 days ago
    Auto-check passed
  • Areno Model Adaptation

    inclusionAI/AReno

    Add or debug an AReno model family, including config conversion, module construction, checkpoint load/save, text or multimodal inference, training backward, tensor parallelism, CUDA graph decode…

    323 GitHub stars~722 tokensUpdated 13 days ago
    Auto-check passed

Works with

Questions about Areno Develop Kernel

What does Areno Develop Kernel do?

Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator. Areno Develop Kernel is an agent skill from inclusionAI/AReno. Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator.

When should I use Areno Develop Kernel?

Areno Develop Kernel fits situations like: changes touch areno/accel; A runtime kernel boundary and require forward; benchmark validation.

How do I install Areno Develop Kernel in Claude Code?

Run `npx skills add inclusionAI/AReno --skill areno-develop-kernel -a claude-code`. Or copy the skill folder (.agents/skills/areno-develop-kernel in inclusionAI/AReno) into .claude/skills/areno-develop-kernel in your project. Claude Code loads it when a task matches its description.

How do I install Areno Develop Kernel in Codex?

Run `npx skills add inclusionAI/AReno --skill areno-develop-kernel -a codex`. Or copy the skill folder (.agents/skills/areno-develop-kernel in inclusionAI/AReno) into .agents/skills/areno-develop-kernel in your project. Codex loads it when a task matches its description.

Can I use Areno Develop Kernel in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add inclusionAI/AReno --skill areno-develop-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/areno-develop-kernel, .gemini/skills/areno-develop-kernel, .github/skills/areno-develop-kernel and .opencode/skills/areno-develop-kernel in your project.

What does Areno Develop Kernel need to run?

Going by SKILL.md and its folder, Areno Develop Kernel needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3.

Does Areno Develop Kernel access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Areno Develop Kernel safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Areno Develop Kernel use?

Areno Develop Kernel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Areno Develop Kernel use?

About 498 tokens (SKILL.md is roughly 2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 166 tokens, read only when the agent opens those files.

What are the alternatives to Areno Develop Kernel?

Skills that share tags, products or a category with Areno Develop Kernel: Esmfold2 (JimLiu/science-skills, 227 stars), MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars), Benchmark Tune (Mesh-LLM/mesh-llm, 3.5k stars) and Cuda Kernel Optimizer (KernelFlow-ops/cuda-optimized-skill, 212 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Areno Develop Kernel?

inclusionAI (a GitHub organization) maintains it in inclusionAI/AReno, which has 323 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on September 24, 2026.

Source: inclusionAI/AReno on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.