Official agent skill

Vectorization

by dotnet in dotnet/skills

Design, implement, optimize, and review SIMD code in .NET. An agent skill from dotnet/skills.

OfficialMITAuto-check passed

Install Vectorization

skills CLI
$ npx skills add dotnet/skills --skill vectorization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dotnet/skills vectorization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dotnet/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/dotnet-advanced/skills/vectorization .claude/skills/vectorization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vectorization
GitHub stars
5.6k
Used in
1 other repo
Token cost
~3k tokens
SKILL.md length
1,481 words
Files
1
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Design, implement, optimize, and review SIMD code in .NET. An agent skill from dotnet/skills.

  • Works in 5 steps: Use the highest-level API that matches… → Start new explicit SIMD loops with… → Keep platforms consistent. Prefer… → …
  • : vectorizing scalar loops with TensorPrimitives
  • SKILL.md covers Inputs and prerequisites, Core rules, Authoring checklist and Testing checklist, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Vectorization is an agent skill from dotnet/skills, published by the product's own GitHub organization. Design, implement, optimize, and review SIMD code in .NET. USE FOR: vectorizing scalar loops with TensorPrimitives, Vector64/128/256/512, or platform hardware intrinsics; reviewing existing SIMD code, including the generic Vector type, for contract equivalence, tail handling, memory safety, portability, fallbacks, and measured performance. DO NOT USE FOR: performance work unrelated to SIMD or vectorization.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with .NET. The repository describes itself as: Repository for skills to assist AI coding agents with .NET and C. The licence is MIT.

When your agent uses it

  • : vectorizing scalar loops with TensorPrimitives
  • Vector64/128/256/512
  • Platform hardware intrinsics
  • Reviewing existing SIMD code

Example prompts

  • “/vectorization”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Use the highest-level API that matches the contract, then stop. Span and string
  2. Start new explicit SIMD loops with Vector128. It is accelerated across the broadest
  3. Keep platforms consistent. Prefer cross-platform operations on the fixed-width vector types;
  4. Read IsHardwareAccelerated, IsSupported, and Count directly. The JIT treats them as
  5. Prefer operators where they are clear. Parenthesize expressions that mix bitwise and

What it can do on your machine

Read from SKILL.md and the folder at commit 8d670fa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are csharp).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • learn.microsoft.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vectorization loads about 3k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 1,481 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dotnet/skills at commit 8d670fa, republished under its MIT licence (© dotnet). 1,481 words, ~2,984 tokens.

Download SKILL.mdSave it as .claude/skills/vectorization/SKILL.md (or your agent's skills folder).
name
vectorization
description
Design, implement, optimize, and review SIMD code in .NET. USE FOR: vectorizing scalar loops with TensorPrimitives, Vector64/128/256/512, or platform hardware intrinsics; reviewing existing SIMD code, including the generic Vector type, for contract equivalence, tail handling, memory safety, portability, fallbacks, and measured performance. DO NOT USE FOR: performance work unrelated to SIMD or vectorization.
license
MIT

.NET SIMD vectorization

Produce a portable optimization that preserves the scalar contract, remains memory-safe at every length, and earns its complexity with measured results. Read the official SIMD and hardware-intrinsics guidance first and follow its comprehensive implementation templates. In particular, use its self-contained per-width dispatch, dedicated small-input handling, loop, and remainder shapes rather than reducing them to a chain of width checks. This skill supplies the decision rules and validation checks to apply while changing real code.

Inputs and prerequisites

Discover these from the repository before asking the user:

InputRequiredWhat to establish
Scalar implementation and testsYesExisting contract, representative call sites, and supported overlap
Target frameworks and platformsYesAvailable SIMD APIs and architectures that must behave consistently
Build and test workflowYesThe repository's normal commands and how to launch separate test processes
Representative workload or benchmarkFor optimizationTypical input sizes and the baseline to beat

Do not add a package merely because an API exists there. First check the target framework and the project's existing dependency/versioning policy.

Core rules

  1. Use the highest-level API that matches the contract, then stop. Span<T> and string operations, TensorPrimitives, and tensor types already accelerate many operations. LINQ reductions such as Sum, Min, Max, and Average can also accelerate when the source exposes its underlying span. Verify empty-input and floating-point behavior rather than assuming similarly named operations are interchangeable. Once an existing API preserves the contract, use it instead of continuing into handwritten SIMD. Before writing an explicit loop, name the framework APIs considered and why none applies. Fixed-shape System.Numerics types remain appropriate for graphics and similar domains.
  2. Start new explicit SIMD loops with Vector128<T>. It is accelerated across the broadest hardware set. Add wider fixed-width paths only when measurements justify them.
  3. Keep platforms consistent. Prefer cross-platform operations on the fixed-width vector types; they lower to the appropriate target instructions. For example, (vector & mask) == Vector128<byte>.Zero becomes ptest on x86/x64. Use architecture-specific intrinsics only for a measured gap, guard them with IsSupported, and retain equivalent portable or scalar behavior.
  4. Read IsHardwareAccelerated, IsSupported, and Count directly. The JIT treats them as constants, so caching them adds no value and obscures which branches disappear.
  5. Prefer operators where they are clear. Parenthesize expressions that mix bitwise and comparison operators so precedence is explicit.

If the task is review-only, do not rewrite the code. Report correctness and memory-safety defects before performance opportunities.

Authoring checklist

  • Contract: identify behavior for empty and short inputs, overlap, overflow, NaN, signed zero, ordering, and exceptions before changing the implementation.
  • Framework gate: inspect the target framework and existing package references, then compile or probe the highest-level candidate API with the required edge cases. A small contract adapter, such as preserving special empty-input behavior, does not justify reimplementing the operation. If the API preserves the contract, use it and stop; do not claim it is unavailable without checking.
  • Structure: for new explicit SIMD, implement Vector128<T> and scalar first. Only after measurements justify wider paths, check Vector512<T>, then Vector256<T>, optional Vector<T>, Vector128<T>, and finally scalar. Omit paths the implementation does not need. Each outer fixed-width guard checks only its IsHardwareAccelerated property and, for generic element types, IsSupported. Inside that block, run the width-specific helper when the input has at least Count elements; otherwise run a dedicated small-input helper, then return. Do not put the length check in the outer guard and fall through to repeat dispatch at narrower widths. Keeping each supported-width block self-contained lets the JIT remove unsupported blocks and avoids redundant work on common small inputs.
  • Loads and stores: prefer span-based Vector128.Create(span) and CopyTo; the JIT keeps them efficient and they require no pinning or reference arithmetic. Unsafe loads and stores are largely unnecessary. When a path genuinely must walk a buffer by managed reference, use the element-offset LoadUnsafe(ref T, nuint) and StoreUnsafe overloads rather than pointers or manually advanced references.
  • Empty inputs: in a reference-based path, obtain the starting reference with MemoryMarshal.GetReference(span) or MemoryMarshal.GetArrayDataReference(array), not by indexing element 0.
  • Unsupported element types: the fixed-width vectors support primitive numeric element types, not char or bool. Reinterpret with MemoryMarshal.Cast or As<TFrom, TTo>; reinterpretation changes only the type, not the bits. Keep Boolean data as 0 or 1 and characters as valid UTF-16, normalizing results before storing when necessary.
  • Offsets: prove the input contains a full vector before subtracting Count or converting an index to nuint; otherwise a negative value becomes a huge unsigned offset.
  • Managed references: do not form references before the start or past the end of a span, including a one-past-end reference. The runtime permits a non-dereferenced managed pointer exactly one past an object or array, but this guidance intentionally prohibits the pattern because it is fragile and easy to misuse. Keep the base reference in range and express traversal with an element offset.
  • Remainders: cover every length, including 0, Count - 1, Count, Count + 1, and nonmultiples of each width. Once the input contains a full vector, keep the tail vectorized by reprocessing the last full vector. An idempotent operation can fold that overlap in directly. A non-idempotent operation must use ConditionalSelect to replace repeated lanes with the operation's identity before folding them in. This is the JIT-recognized general pattern; it can reduce a zero-identity selection to a bitwise mask while retaining broader optimization opportunities. For in-place transforms, preserve the original tail values before overlapping stores and write only valid results.
  • Buffer overlap: choose a traversal direction or staging strategy that prevents stores from corrupting values not yet loaded.
  • Numeric behavior: account for floating-point reassociation, NaN and signed-zero semantics, checked or unchecked integer overflow, and endianness where the algorithm depends on byte order. Native and Estimate operations can intentionally relax precision or IEEE edge-case behavior; use them only when the contract permits it and measurements justify them.

The official guidance contains the complete dispatch, small-input, unrolling, and remainder templates; use those for the full implementation. The following excerpt illustrates only the inner safe Vector128<T> loop for an in-place elementwise transform, after its self-contained dispatch block has established at least one full vector. Transform represents the operation being implemented:

csharp
Span<int> tail = data.Slice(data.Length - Vector128<int>.Count);
Vector128<int> end = Vector128.Create<int>(tail);
Span<int> remaining = data;

while (remaining.Length >= Vector128<int>.Count)
{
    Vector128<int> values = Vector128.Create<int>(remaining);
    Transform(values).CopyTo(remaining);
    remaining = remaining.Slice(Vector128<int>.Count);
}

if (!remaining.IsEmpty)
{
    Transform(end).CopyTo(tail);
}

The early end load preserves original values before overlapping stores. For a read-only reduction, load the same final span after the main loop and use ConditionalSelect to replace already-processed lanes with the operation's identity. Do not substitute LoadUnsafe/StoreUnsafe or a scalar epilogue merely to avoid span bounds checks.

Show full SKILL.md (420 more words)Show less

Testing checklist

  • Compare the optimized implementation with the scalar contract across boundary lengths, randomized values, empty inputs, supported overlap, and numeric edge cases. Cover every implemented width and the scalar path with inputs both large enough and too small to benefit.
  • Exercise every implemented width and the scalar fallback in separate processes. On x86/x64 CoreCLR, DOTNET_EnableAVX2=0 disables AVX2 and DOTNET_EnableHWIntrinsic=0 disables hardware intrinsics. Use the repository's normal test command and do not change these process-wide settings inside a unit test. These settings do not change code already compiled as ReadyToRun or ahead of time, so confirm the target code is JIT-compiled when using them to force a path.
  • For unsafe loads and stores, use guard-page or equivalent boundary tests when available. Put the inaccessible page after the buffer for forward iteration and before it for backwards iteration, and include nonmultiple lengths. An ordinary array allocation does not reliably expose an out-of-bounds read.

Benchmarking

Use BenchmarkDotNet to measure representative small and large inputs before keeping the added complexity. Compare scalar, Vector128<T>, and each wider implemented path in the same run. Small inputs can be slower because setup dominates, and speedups are rarely the theoretical vector-width multiple because memory throughput, alignment, and latency still apply. Report throughput or time with noise context and, when relevant, generated code size or instruction counts. Control allocation alignment for stable measurements or randomize it to observe the distribution. A wider vector is not automatically faster.

If the project cannot target the required framework, run the relevant architecture, or execute the fallback configuration, state exactly which path remains unverified. Do not claim success from a default-hardware test alone.

Completion contract

  • Authoring: leave the scalar contract covered by tests; identify the framework or SIMD layer selected; report measurements for the representative workload; name any architecture or fallback path that could not be exercised.
  • Review: report only concrete findings, ordered by correctness, memory safety, portability, tests, then performance evidence. If none remain, say so directly.
  • Do not call an optimization complete when it only builds, only passes on the current machine, or has no comparison against the scalar baseline.

Review checklist

Review in this order:

  1. Scalar-contract equivalence, including signed zero, NaN, overflow, and relevant endianness
  2. Reuse of an existing accelerated framework API
  3. Tail correctness for idempotent versus non-idempotent work
  4. Memory safety, unsigned offset arithmetic, empty inputs, and overlapping buffers
  5. Portable dispatch and behaviorally equivalent fallbacks
  6. Tests that force each width and the scalar path
  7. Benchmarks that justify explicit SIMD and additional widths

© dotnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/dotnet-advanced/skills/vectorization of dotnet/skills.

Open the folder on GitHubat commit 8d670fa

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in dotnet/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Vectorization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vectorization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vectorization this skilldotnet/skills5.6k1 repos~3kAutomated safety check: PassMIT
Minimax DOCXpoco-ai/poco-claw1.4k7 repos~3.9kAutomated safety check: PassMIT
Microsoft Skill CreatorMicrosoftDocs/mcp1.9k3 repos~2.1kAutomated safety check: PassCC-BY-4.0
Speckit ConstitutionWeihanLi/WeihanLi.Common24211 repos~2.1kAutomated safety check: PassApache-2.0
Copilot Session Failure Analysisdotnet/maui23k—~3.4kAutomated safety check: PassMIT
Microsoft Code ReferenceMicrosoftDocs/mcp1.9k4 repos~1.1kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Minimax DOCX

    poco-ai/poco-claw

    Professional DOCX document creation, editing, and formatting using OpenXML SDK (.NET).

    1.4k GitHub starsUsed in 7 repos~3.9k tokens
    Documents & OfficeAuto-check passed
  • Microsoft Skill Creator

    MicrosoftDocs/mcp

    Official

    Create agent skills for Microsoft technologies using official documentation.

    1.9k GitHub starsUsed in 3 repos~2.1k tokens
    Agent WorkflowsAuto-check passed
  • Speckit Constitution

    WeihanLi/WeihanLi.Common

    Create or update the project constitution from interactive or provided principle inputs, ensuring all dependent templates stay in sync.

    242 GitHub starsUsed in 11 repos~2.1k tokens
    DevelopmentAuto-check passed
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Microsoft Code Reference

    MicrosoftDocs/mcp

    Official

    Find working code samples, verify API signatures, and fix Microsoft SDK errors using official docs.

    1.9k GitHub starsUsed in 4 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Official

    Audits and updates os-packages.json files listing the Linux packages each .NET release needs per distro, then regenerates the Markdown from the JSON.

    22k GitHub stars~2.3k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from dotnet/skills

All 91 skills in this repo
  • Official

    Resolves .NET runtime frames in Apple .ips crash logs to function names, source files and line numbers using dSYM symbols, atos and the Microsoft symbol server.

    5.6k GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Official

    Resolves native crash frames from .NET Android tombstones to function names, source files and line numbers using BuildIds, Microsoft's symbol server and llvm-symbolizer.

    5.6k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Official

    Scans C# and .NET code for about 50 performance anti-patterns and reports prioritized findings with concrete fixes, at a scan depth you choose.

    5.6k GitHub starsUsed in 3 repos~3.1k tokens
    Auto-check passed
  • Official

    Statically pairs source files with test files to list code that no test references, using Roslyn for C# or tree-sitter for many languages, with no build.

    5.6k GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check passed
  • Microbenchmarking

    dotnet/skills

    Official

    Activate this skill when BenchmarkDotNet (BDN) is involved in the task — creating, running, configuring, or reviewing BDN benchmarks.

    5.6k GitHub starsUsed in 3 repos~3.3k tokens
    Auto-check passed
  • Official

    Makes .NET projects compatible with Native AOT and trimming by resolving IL trim and AOT analyzer warnings through annotations rather than suppressions.

    5.6k GitHub starsUsed in 2 repos~4.2k tokens
    Auto-check passed

Works with

Questions about Vectorization

What does Vectorization do?

Design, implement, optimize, and review SIMD code in .NET. An agent skill from dotnet/skills. Vectorization is an agent skill from dotnet/skills, published by the product's own GitHub organization.NET.

When should I use Vectorization?

Vectorization fits situations like: : vectorizing scalar loops with TensorPrimitives; vector64/128/256/512; platform hardware intrinsics; reviewing existing SIMD code.

How do I install Vectorization in Claude Code?

Run `npx skills add dotnet/skills --skill vectorization -a claude-code`. Or copy the skill folder (plugins/dotnet-advanced/skills/vectorization in dotnet/skills) into .claude/skills/vectorization in your project. Claude Code loads it when a task matches its description.

How do I install Vectorization in Codex?

Run `npx skills add dotnet/skills --skill vectorization -a codex`. Or copy the skill folder (plugins/dotnet-advanced/skills/vectorization in dotnet/skills) into .agents/skills/vectorization in your project. Codex loads it when a task matches its description.

Can I use Vectorization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dotnet/skills --skill vectorization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vectorization, .gemini/skills/vectorization, .github/skills/vectorization and .opencode/skills/vectorization in your project.

What does Vectorization need to run?

SKILL.md names no scripts, command-line tools or credentials: Vectorization is instructions for the agent only.

Does Vectorization access the network?

SKILL.md names 1 domain. As links in the text: learn.microsoft.com. This is read from the text; nothing was executed.

Is Vectorization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vectorization use?

Vectorization is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vectorization use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vectorization?

Skills that share tags, products or a category with Vectorization: Minimax DOCX (poco-ai/poco-claw, 1.4k stars), Microsoft Skill Creator (MicrosoftDocs/mcp, 1.9k stars), Speckit Constitution (WeihanLi/WeihanLi.Common, 242 stars) and Copilot Session Failure Analysis (dotnet/maui, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vectorization?

dotnet (a GitHub organization, an official publisher) maintains it in dotnet/skills, which has 5,568 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on October 7, 2026.

Source: dotnet/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.