Agent skill

Transformer Architecture Guide

by wentorai in wentorai/research-plugins

Guide to Transformer architectures for NLP and computer vision

MITAuto-check passedAI & LLM Engineering

Install Transformer Architecture Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill transformer-architecture-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins transformer-architecture-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/domains/ai-ml/transformer-architecture-guide .claude/skills/transformer-architecture-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
transformer-architecture-guide
GitHub stars
298
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
323 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Guide to Transformer architectures for NLP and computer vision

  • Tasks that involve Computer vision
  • SKILL.md covers The Original Transformer, Major Transformer Variants, Vision Transformers (ViT) and Efficient Transformer Variants, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Natural language processing

What it does

Transformer Architecture Guide is an agent skill from wentorai/research-plugins. Guide to Transformer architectures for NLP and computer vision

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Computer vision and Natural language processing. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Computer vision
  • Tasks that involve Natural language processing

Example prompts

  • “/transformer-architecture-guide”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Transformer Architecture Guide loads about 2.1k tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 323 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 323 words, ~2,109 tokens.

Download SKILL.mdSave it as .claude/skills/transformer-architecture-guide/SKILL.md (or your agent's skills folder).
name
transformer-architecture-guide
description
Guide to Transformer architectures for NLP and computer vision

Transformer Architecture Guide

Understand, implement, and adapt Transformer architectures for NLP, computer vision, and multimodal research, from the original attention mechanism to modern variants.

The Original Transformer

The Transformer (Vaswani et al., 2017, "Attention Is All You Need") replaced recurrence and convolution with self-attention as the primary sequence modeling mechanism.

Core Components
ComponentFunctionKey Parameters
Multi-Head Self-AttentionComputes attention weights across all positionsd_model, n_heads, d_k, d_v
Feed-Forward NetworkPosition-wise nonlinear transformationd_model, d_ff
Positional EncodingInjects sequence order informationSinusoidal or learned
Layer NormalizationStabilizes trainingPre-norm or post-norm
Residual ConnectionsEnables gradient flow in deep networksAdd before or after norm
Self-Attention Mechanism
python
import torch
import torch.nn as nn
import torch.nn.functional as F
import math

class MultiHeadAttention(nn.Module):
    def __init__(self, d_model=512, n_heads=8):
        super().__init__()
        self.d_model = d_model
        self.n_heads = n_heads
        self.d_k = d_model // n_heads

        self.W_q = nn.Linear(d_model, d_model)
        self.W_k = nn.Linear(d_model, d_model)
        self.W_v = nn.Linear(d_model, d_model)
        self.W_o = nn.Linear(d_model, d_model)

    def forward(self, Q, K, V, mask=None):
        batch_size = Q.size(0)

        # Linear projections and reshape for multi-head
        Q = self.W_q(Q).view(batch_size, -1, self.n_heads, self.d_k).transpose(1, 2)
        K = self.W_k(K).view(batch_size, -1, self.n_heads, self.d_k).transpose(1, 2)
        V = self.W_v(V).view(batch_size, -1, self.n_heads, self.d_k).transpose(1, 2)

        # Scaled dot-product attention
        scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(self.d_k)
        if mask is not None:
            scores = scores.masked_fill(mask == 0, -1e9)
        attn_weights = F.softmax(scores, dim=-1)
        context = torch.matmul(attn_weights, V)

        # Concatenate heads and project
        context = context.transpose(1, 2).contiguous().view(batch_size, -1, self.d_model)
        return self.W_o(context)
Complete Transformer Block
python
class TransformerBlock(nn.Module):
    def __init__(self, d_model=512, n_heads=8, d_ff=2048, dropout=0.1):
        super().__init__()
        self.attention = MultiHeadAttention(d_model, n_heads)
        self.norm1 = nn.LayerNorm(d_model)
        self.norm2 = nn.LayerNorm(d_model)
        self.ffn = nn.Sequential(
            nn.Linear(d_model, d_ff),
            nn.GELU(),
            nn.Dropout(dropout),
            nn.Linear(d_ff, d_model),
            nn.Dropout(dropout)
        )
        self.dropout = nn.Dropout(dropout)

    def forward(self, x, mask=None):
        # Pre-norm architecture (GPT-style)
        attn_out = self.attention(self.norm1(x), self.norm1(x), self.norm1(x), mask)
        x = x + self.dropout(attn_out)
        ffn_out = self.ffn(self.norm2(x))
        x = x + ffn_out
        return x

Major Transformer Variants

Architecture Taxonomy
ArchitectureTypeKey InnovationRepresentative Model
Encoder-onlyBidirectionalMasked language modelingBERT, RoBERTa
Decoder-onlyAutoregressiveCausal language modelingGPT, LLaMA, Claude
Encoder-DecoderSeq2seqCross-attention between encoder and decoderT5, BART, mBART
Encoder-Only (BERT Family)
python
# BERT-style masked language modeling
from transformers import BertTokenizer, BertForMaskedLM

tokenizer = BertTokenizer.from_pretrained("bert-base-uncased")
model = BertForMaskedLM.from_pretrained("bert-base-uncased")

text = "The Transformer architecture has [MASK] natural language processing."
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)

# Get predictions for [MASK]
mask_idx = (inputs.input_ids == tokenizer.mask_token_id).nonzero(as_tuple=True)[1]
logits = outputs.logits[0, mask_idx]
top_tokens = logits.topk(5).indices[0]
print([tokenizer.decode(t) for t in top_tokens])
Decoder-Only (GPT Family)
python
# GPT-style autoregressive generation
from transformers import GPT2LMHeadModel, GPT2Tokenizer

tokenizer = GPT2Tokenizer.from_pretrained("gpt2")
model = GPT2LMHeadModel.from_pretrained("gpt2")

prompt = "The key innovation of the Transformer is"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(
    **inputs,
    max_new_tokens=50,
    temperature=0.7,
    top_p=0.9,
    do_sample=True
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Vision Transformers (ViT)

The Vision Transformer (Dosovitskiy et al., 2021) applies the Transformer to image classification:

python
class VisionTransformer(nn.Module):
    def __init__(self, img_size=224, patch_size=16, in_channels=3,
                 d_model=768, n_heads=12, n_layers=12, n_classes=1000):
        super().__init__()
        self.patch_size = patch_size
        n_patches = (img_size // patch_size) ** 2

        # Patch embedding: split image into patches and project
        self.patch_embed = nn.Conv2d(in_channels, d_model,
                                     kernel_size=patch_size, stride=patch_size)

        # Learnable [CLS] token and position embeddings
        self.cls_token = nn.Parameter(torch.zeros(1, 1, d_model))
        self.pos_embed = nn.Parameter(torch.zeros(1, n_patches + 1, d_model))

        # Transformer blocks
        self.blocks = nn.ModuleList([
            TransformerBlock(d_model, n_heads) for _ in range(n_layers)
        ])

        self.norm = nn.LayerNorm(d_model)
        self.head = nn.Linear(d_model, n_classes)

    def forward(self, x):
        B = x.size(0)
        # Patchify and flatten
        x = self.patch_embed(x).flatten(2).transpose(1, 2)  # (B, n_patches, d_model)

        # Prepend CLS token
        cls = self.cls_token.expand(B, -1, -1)
        x = torch.cat([cls, x], dim=1)
        x = x + self.pos_embed

        # Transformer blocks
        for block in self.blocks:
            x = block(x)

        # Classification from CLS token
        x = self.norm(x[:, 0])
        return self.head(x)

Efficient Transformer Variants

MethodComplexityKey IdeaReference
Standard attentionO(n^2)Full pairwise attentionVaswani et al., 2017
Linear attentionO(n)Kernel approximation of softmaxKatharopoulos et al., 2020
Flash AttentionO(n^2) time, O(n) memoryIO-aware tiled computationDao et al., 2022
Sparse attentionO(n sqrt(n))Fixed or learned sparse patternsChild et al., 2019
Sliding windowO(n * w)Local attention windowBeltagy et al., 2020 (Longformer)
Multi-query attentionO(n^2) but fasterShared K/V across headsShazeer, 2019
Grouped-query attentionO(n^2) but fasterGroups of heads share K/VAinslie et al., 2023

Model Scaling Laws

Kaplan et al. (2020) and Hoffmann et al. (2022, "Chinchilla") established scaling laws:

Performance (loss) scales as a power law with:
- Model parameters (N): L ~ N^(-0.076)
- Dataset size (D): L ~ D^(-0.095)
- Compute budget (C): L ~ C^(-0.050)

Chinchilla optimal scaling:
- For compute budget C, allocate equally to model size and data
- Optimal tokens ~ 20 * parameters
- Example: 70B parameter model needs ~1.4T training tokens

Research Resources

ResourceDescription
Hugging Face TransformersPre-trained models and fine-tuning framework
Papers With CodeBenchmarks, SOTA tracking, and code links
The Illustrated Transformer (Jay Alammar)Visual explanations of attention
Andrej Karpathy's nanoGPTMinimal GPT implementation for education
EleutherAIOpen-source LLM research community
MLCommonsStandardized ML benchmarks (MLPerf)

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/domains/ai-ml/transformer-architecture-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Transformer Architecture Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Transformer Architecture Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Transformer Architecture Guide this skillwentorai/research-plugins2981 repos~2.1kAutomated safety check: PassMIT
Hugging Face Transformers Usagedavila7/claude-code-templates33k11 repos~1.2kAutomated safety check: PassMIT
Transformersynulihao/AgentSkillOS618—~2.9kAutomated safety check: PassNone
Scholar Computejoshzyj/open-scholar-skill168—~15kAutomated safety check: PassCustom licence
Deep Learning NLPDrchronx/ai-agent-research-starter-kit139—~516Automated safety check: PassCustom licence
Transformers.jshuggingface/skills11k1 repos~6.2kAutomated safety check: PassApache-2.0

Similar skills

  • Hugging Face Transformers Usage

    davila7/claude-code-templates

    Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.

    33k GitHub starsUsed in 11 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Transformers

    ynulihao/AgentSkillOS

    Work with state-of-the-art machine learning models for NLP, computer vision, audio, and multimodal tasks using HuggingFace Transformers.

    618 GitHub stars~2.9k tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • Scholar Compute

    joshzyj/open-scholar-skill

    Design and execute computational social science analyses across 11 modules: text-as-data/NLP (STM, BERTopic, Wordfish, BERT, conText embedding regression, LLM annotation + DSL bias correction…

    168 GitHub stars~15k tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • Deep Learning NLP

    Drchronx/ai-agent-research-starter-kit

    Paddle-based deep learning workflows from the course materials, including DNN/RNN text-style baselines and the CNN/LeNet image classification case using folder-labeled digit images.

    139 GitHub stars~516 tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Transformers.js

    huggingface/skills

    Official

    Runs pre-trained Hugging Face models in JavaScript or TypeScript with Transformers.js, in browsers or Node.js, Bun and Deno, for text, vision, audio and multimodal tasks.

    11k GitHub starsUsed in 1 repo~6.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Find AI Consultancy

    jeremylongshore/tons-of-skills-marketplace

    A skill your agent uses whenever the user wants to find, shortlist, vet, or enrich US AI/ML/data consulting firms (consultancies) — AI/ML development, MLOps, generative AI / LLM apps (RAG, chatbots…

    2.8k GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Transformer Architecture Guide

What does Transformer Architecture Guide do?

Guide to Transformer architectures for NLP and computer vision. Transformer Architecture Guide is an agent skill from wentorai/research-plugins.

When should I use Transformer Architecture Guide?

Transformer Architecture Guide fits situations like: tasks that involve Computer vision; tasks that involve Natural language processing.

How do I install Transformer Architecture Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill transformer-architecture-guide -a claude-code`. Or copy the skill folder (skills/domains/ai-ml/transformer-architecture-guide in wentorai/research-plugins) into .claude/skills/transformer-architecture-guide in your project. Claude Code loads it when a task matches its description.

How do I install Transformer Architecture Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill transformer-architecture-guide -a codex`. Or copy the skill folder (skills/domains/ai-ml/transformer-architecture-guide in wentorai/research-plugins) into .agents/skills/transformer-architecture-guide in your project. Codex loads it when a task matches its description.

Can I use Transformer Architecture Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill transformer-architecture-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/transformer-architecture-guide, .gemini/skills/transformer-architecture-guide, .github/skills/transformer-architecture-guide and .opencode/skills/transformer-architecture-guide in your project.

What does Transformer Architecture Guide need to run?

SKILL.md names no scripts, command-line tools or credentials: Transformer Architecture Guide is instructions for the agent only. Our summary lists: Python 3.

Does Transformer Architecture Guide access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Transformer Architecture Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Transformer Architecture Guide use?

Transformer Architecture Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Transformer Architecture Guide use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Transformer Architecture Guide?

Skills that share tags, products or a category with Transformer Architecture Guide: Hugging Face Transformers Usage (davila7/claude-code-templates, 33k stars), Transformers (ynulihao/AgentSkillOS, 618 stars), Scholar Compute (joshzyj/open-scholar-skill, 168 stars) and Deep Learning NLP (Drchronx/ai-agent-research-starter-kit, 139 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Transformer Architecture Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.