Agent skill

Attention Variants From Papers

by benchflow-ai in benchflow-ai/skillsbench

Implement paper-defined attention variants by extracting mechanism invariants, mapping tensor shapes, preserving module interfaces, validating numerical behavior, and integrating the custom…

Apache-2.0Auto-check passed

Install Attention Variants From Papers

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill attention-variants-from-papers -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench attention-variants-from-papers --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks-extra/diff-transformer_impl/environment/skills/attention-variants-from-papers .claude/skills/attention-variants-from-papers && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
attention-variants-from-papers
GitHub stars
1.8k
Token cost
~669 tokens
SKILL.md length
260 words
Files
6 (incl. references)
Skills in repo
189
Repo updated
First seen
Licence
Apache-2.0

At a glance

Implement paper-defined attention variants by extracting mechanism invariants, mapping tensor shapes, preserving module interfaces, validating numerical behavior, and integrating the custom…

  • Works in 6 steps: Read the paper for invariants, not names. → Build a shape ledger before coding. → Choose the mechanism pattern. → …
  • SKILL.md covers Workflow, Checklist and Reference Map
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Attention Variants From Papers is an agent skill from benchflow-ai/skillsbench. Implement paper-defined attention variants by extracting mechanism invariants, mapping tensor shapes, preserving module interfaces, validating numerical behavior, and integrating the custom attention into an existing transformer stack.

Its SKILL.md is about 670 tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/mechanism-patterns.md`, `references/paper-to-implementation.md` and `references/shape-ledger.md`).

The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

Example prompts

  • “/attention-variants-from-papers”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Read the paper for invariants, not names.
  2. Build a shape ledger before coding.
  3. Choose the mechanism pattern.
  4. Preserve the module boundary.
  5. Validate in layers.
  6. Integrate into the stack last.

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Attention Variants From Papers loads about 669 tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 260 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~669
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 260 words, ~669 tokens.

Download SKILL.mdSave it as .claude/skills/attention-variants-from-papers/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
attention-variants-from-papers
description
Implement paper-defined attention variants by extracting mechanism invariants, mapping tensor shapes, preserving module interfaces, validating numerical behavior, and integrating the custom attention into an existing transformer stack.

Attention Variants from Papers

Use this skill when a paper changes how attention scores, branches, normalization, or head sharing work, but the module still needs to behave like a drop-in transformer attention block.

Workflow

  1. Read the paper for invariants, not names. Use paper-to-implementation.md to extract the external contract, the changed computation, and the training-time constraints.
  2. Build a shape ledger before coding. Use shape-ledger.md to track projections, head grouping, branch count, and output width.
  3. Choose the mechanism pattern. Use mechanism-patterns.md for subtractive attention, branch mixing, learned gates, and extra normalization.
  4. Preserve the module boundary. Keep the same input and output shape, mask semantics, positional encoding flow, and cache behavior unless the task explicitly changes them.
  5. Validate in layers. Start with random-tensor smoke tests, then compare against a baseline attention path. Use stability-and-validation.md.
  6. Integrate into the stack last. Swap the new module into one transformer block, verify the residual path, then roll it through the full model. Use transformer-integration.md.

Checklist

  • extract the paper's invariants before writing code
  • account for every reshape, branch, and repeat in a shape ledger
  • preserve output width at concatenation or output projection
  • apply masks and positional terms at the intended stage
  • confirm random smoke tests stay finite
  • compare unchanged behaviors against a baseline attention implementation

Reference Map

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in tasks-extra/diff-transformer_impl/environment/skills/attention-variants-from-papers of benchflow-ai/skillsbench.

  • SKILL.md
  • references/mechanism-patterns.md
  • references/paper-to-implementation.md
  • references/shape-ledger.md
  • references/stability-and-validation.md
  • references/transformer-integration.md

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Attention Variants From Papers next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Attention Variants From Papers compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Attention Variants From Papers this skillbenchflow-ai/skillsbench1.8k—~669Automated safety check: PassApache-2.0
Implementsickn33/agentic-awesome-skills47k5 repos~306Automated safety check: PassMIT
Implementcodewhale-hq/Codewhale41k—~190Automated safety check: PassMIT
Papers Skillsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
Bgpt Paper SearchK-Dense-AI/scientific-agent-skills48k1 repos~2kAutomated safety check: PassMIT
Extractalirezarezvani/claude-skills28k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Implement

    sickn33/agentic-awesome-skills

    Implement a piece of work based on a PRD or set of issues. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 5 repos~306 tokens
    Product & Project ManagementAuto-check passed
  • Implement

    codewhale-hq/Codewhale

    Carry an authorized, defined request or approved plan through scoped edits and proportionate verification.

    41k GitHub stars~190 tokensUpdated today
    Auto-check passed
  • Papers Skill

    sickn33/agentic-awesome-skills

    Skill for academic research workflows: search Semantic Scholar (200M+ papers), inspect citations, download arXiv PDFs, and extract PDF text.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Research & ScienceAuto-check passed
  • Bgpt Paper Search

    K-Dense-AI/scientific-agent-skills

    Searches BGPT scientific papers by topic or DOI and retrieves claim-level evidence extracted from full text, including experiments, reported statistics, scope, limitations, and provenance.

    48k GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Extract

    alirezarezvani/claude-skills

    Turn a proven pattern or debugging solution into a standalone reusable skill with SKILL.md, reference docs, and examples.

    28k GitHub stars~1.4k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Slo Implementation

    wshobson/agents

    Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting.

    40k GitHub starsUsed in 11 repos~1.7k tokens
    DevOps & CloudAuto-check passed

More from benchflow-ai/skillsbench

All 189 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Attention Variants From Papers

What does Attention Variants From Papers do?

Implement paper-defined attention variants by extracting mechanism invariants, mapping tensor shapes, preserving module interfaces, validating numerical behavior, and integrating the custom…. Attention Variants From Papers is an agent skill from benchflow-ai/skillsbench. Implement paper-defined attention variants by extracting mechanism invariants, mapping tensor shapes, preserving module interfaces, validating numerical behavior, and integrating the custom attention into an existing transformer stack.

How do I install Attention Variants From Papers in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill attention-variants-from-papers -a claude-code`. Or copy the skill folder (tasks-extra/diff-transformer_impl/environment/skills/attention-variants-from-papers in benchflow-ai/skillsbench) into .claude/skills/attention-variants-from-papers in your project. Claude Code loads it when a task matches its description.

How do I install Attention Variants From Papers in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill attention-variants-from-papers -a codex`. Or copy the skill folder (tasks-extra/diff-transformer_impl/environment/skills/attention-variants-from-papers in benchflow-ai/skillsbench) into .agents/skills/attention-variants-from-papers in your project. Codex loads it when a task matches its description.

Can I use Attention Variants From Papers in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill attention-variants-from-papers -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/attention-variants-from-papers, .gemini/skills/attention-variants-from-papers, .github/skills/attention-variants-from-papers and .opencode/skills/attention-variants-from-papers in your project.

What does Attention Variants From Papers need to run?

SKILL.md names no scripts, command-line tools or credentials: Attention Variants From Papers is instructions for the agent only.

Does Attention Variants From Papers access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Attention Variants From Papers safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Attention Variants From Papers use?

Attention Variants From Papers is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Attention Variants From Papers use?

About 669 tokens (SKILL.md is roughly 2.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.

What are the alternatives to Attention Variants From Papers?

Skills that share tags, products or a category with Attention Variants From Papers: Implement (sickn33/agentic-awesome-skills, 47k stars), Implement (codewhale-hq/Codewhale, 41k stars), Papers Skill (sickn33/agentic-awesome-skills, 47k stars) and Bgpt Paper Search (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Attention Variants From Papers?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,835 GitHub stars. The repository holds 189 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.