JS Perf Investigation
SAP/project-foxhound
Structured performance opportunity investigation for SpiderMonkey (the Firefox JavaScript engine).
Guidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo.
$ npx skills add fla-org/flash-linear-attention --skill fla-ascend-performance -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install fla-org/flash-linear-attention fla-ascend-performance --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/fla-org/flash-linear-attention.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/fla-ascend-performance .claude/skills/fla-ascend-performance && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "fla-ascend-performance" agent skill from https://github.com/fla-org/flash-linear-attention/tree/main/.agents/skills/fla-ascend-performance into .claude/skills/fla-ascend-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fla-ascend-performance", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/fla-org/flash-linear-attention/tree/main/.agents/skills/fla-ascend-performanceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add fla-org/flash-linear-attention --skill fla-ascend-performance -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install fla-org/flash-linear-attention fla-ascend-performance --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/fla-org/flash-linear-attention.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/fla-ascend-performance .agents/skills/fla-ascend-performance && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "fla-ascend-performance" agent skill from https://github.com/fla-org/flash-linear-attention/tree/main/.agents/skills/fla-ascend-performance into .agents/skills/fla-ascend-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fla-ascend-performance", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add fla-org/flash-linear-attention --skill fla-ascend-performance -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install fla-org/flash-linear-attention fla-ascend-performance --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/fla-org/flash-linear-attention.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/fla-ascend-performance .cursor/skills/fla-ascend-performance && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "fla-ascend-performance" agent skill from https://github.com/fla-org/flash-linear-attention/tree/main/.agents/skills/fla-ascend-performance into .cursor/skills/fla-ascend-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fla-ascend-performance", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/fla-org/flash-linear-attention.git --path .agents/skills/fla-ascend-performance--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add fla-org/flash-linear-attention --skill fla-ascend-performance -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install fla-org/flash-linear-attention fla-ascend-performance --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/fla-org/flash-linear-attention.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/fla-ascend-performance .gemini/skills/fla-ascend-performance && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "fla-ascend-performance" agent skill from https://github.com/fla-org/flash-linear-attention/tree/main/.agents/skills/fla-ascend-performance into .gemini/skills/fla-ascend-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fla-ascend-performance", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install fla-org/flash-linear-attention fla-ascend-performanceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add fla-org/flash-linear-attention --skill fla-ascend-performance -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/fla-org/flash-linear-attention.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/fla-ascend-performance .github/skills/fla-ascend-performance && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "fla-ascend-performance" agent skill from https://github.com/fla-org/flash-linear-attention/tree/main/.agents/skills/fla-ascend-performance into .github/skills/fla-ascend-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fla-ascend-performance", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add fla-org/flash-linear-attention --skill fla-ascend-performance -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install fla-org/flash-linear-attention fla-ascend-performance --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/fla-org/flash-linear-attention.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/fla-ascend-performance .opencode/skills/fla-ascend-performance && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "fla-ascend-performance" agent skill from https://github.com/fla-org/flash-linear-attention/tree/main/.agents/skills/fla-ascend-performance into .opencode/skills/fla-ascend-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fla-ascend-performance", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
fla-ascend-performanceGuidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo.
Fla Ascend Performance is an agent skill from fla-org/flash-linear-attention. Guidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo. Covers profiling with torchnpu, PipeUtilization/MemoryUB CSV analysis, Cube/Vector/MTE/UB bottleneck diagnosis, and kernel optimization (UB tiling, grid splits, fusion/split, varlen, GTCONTIG gate loading, constexpr DMA-path split / TAILMODE, extractslice, MTE OOB, int32 address overflow, tl.cast vs constexpr .to, makeblockptr int32 offsets, correctness gates, tl.dot left-operand clobber). NPU kernels must not use…
Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/TRAPS.md`, `references/cases.md` and `references/g-contiguous-loading.md`).
It sits in Security, covering Performance optimization, Statistics and Threat modeling. The repository describes itself as: 🚀 Efficient implementations for emerging model architectures. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b8ff848. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonrgFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Fla Ascend Performance loads about 5.6k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 219 tokens; SKILL.md has 2,492 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from fla-org/flash-linear-attention at commit b8ff848, republished under its MIT licence (© fla-org). 2,492 words, ~5,596 tokens.
.claude/skills/fla-ascend-performance/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.Use this skill for Ascend operator performance work on all files under any triton_ascend directory (**/triton_ascend/**).
Multi-round iteration discipline (frozen tests, task contract, when to stop): fla-optimization-loop. MR packaging: fla-mr-readiness.
Collection must use this skill's generic scripts — do not copy torch_npu.profiler boilerplate per op.
Environment: Use the Python/NPU environment already active in the current terminal (including any activated conda/venv). Run collection, analysis, and benchmarks in the same shell; do not spawn a new shell or switch environments mid-workflow. If the terminal has no NPU stack loaded yet, activate the project's Ascend environment first, then continue in that same session. Metrics, failure modes, code index: reference.md. Past kernel notes: cases.md.
Make the target backend semantically correct before optimizing; never hide missing capability or kernel bugs behind a Torch fallback. What generalizes: UB modeling, grid splits, layout/precision, and verification. Values like BC=16, K slabs of 64, or specific mem_mult are starting points only — do not copy them as rules.
Hard constraint (NPU launch params): Ascend Triton kernels do not support num_warps or num_stages. During optimization these kwargs must never appear in @triton.jit launches, triton.autotune configs — do not copy them from CUDA Triton. Tune via tiles, grid, layout, fusion/split, and UB budget only.
- [ ] 1. Freeze semantics, workload, and baseline latency
- [ ] 2. Collect with generic scripts (first pass: PipeUtilization)
- [ ] 3. Parse CSVs and classify the bottleneck
- [ ] 4. Triton-Ascend optimize for that bottleneck (MemoryUB if needed)
- [ ] 5. Correctness gate + synchronized benchmark
- [ ] 6. Re-profile to confirm metrics, then decide whether to continue@dispatch, default impl, and closest Ascend impl; list layout, dtype, fixed/varlen, head mapping, fwd/bwd, and optional args.torch.npu.synchronize() + repeats); confirm the target NPU kernel runs, not a Torch fallback.IS_NPU with lazy imports; verifiers must state real support ranges.Scripts live under .agents/skills/fla-ascend-performance/scripts/ (run from that directory or set PYTHONPATH).
| Script | Role |
|---|---|
scripts/profile_npu.py | Trace any workload() |
scripts/analyze_profile.py | Parse op_statistic / kernel_details |
SKILL_DIR=.agents/skills/fla-ascend-performance
cd "$SKILL_DIR"
python scripts/profile_npu.py \
--name my_op --out-dir npu_prof \
--metrics PipeUtilization --analyze \
--kernel-filter my_kernel_substr \
--exec-file path/to/workload_only.pyworkload_only.py only defines workload() — no profiler boilerplate:
def workload():
y = op(...)
y.backward(grad)Library usage (when not using --exec-file):
from profile_npu import profile_callable
def workload():
y = op(...)
y.backward(grad)
trace_dir = profile_callable(
workload,
name="my_op",
out_dir="npu_prof",
aic_metrics="PipeUtilization", # or MemoryUB / L2Cache / ...
)Default schedule: wait=0, warmup=1, active=1, repeat=1. One aic_metrics per run; start with PipeUtilization, collect MemoryUB separately for UB bandwidth.
cd .agents/skills/fla-ascend-performance
python scripts/analyze_profile.py path/to/*_profiling_* --kernel-filter <substr>op_statistic: who owns Total Time; is the target kernel the real hotspot?kernel_details (by Duration): read pipe / UB columns.| Signal | Bottleneck | Prefer |
|---|---|---|
High aiv_vec_ratio, Cube≈0 | Vector-bound | Larger row tile, less scalar, fuse load/store |
High aic_mac_ratio / cube_utilization | Cube-bound | Better matmul tiles/alignment, less non-Cube prelude |
High mte2/mte3_ratio, low compute | Memory-move-bound | More reuse, fewer writebacks; check strides — gate g stride-HV gather often 10×+ slower (g-contiguous-loading.md) |
High scalar_ratio | Scalar-bound | Vectorize, kill branches, heuristics |
| High UB bw under MemoryUB, low vec/mac | UB bandwidth saturated | Larger tiles / more fusion |
| Low target Ratio, many tiny ops | Unfused / fallback | Fix dispatch and fusion first |
Two kernels share o + high MTE | Intermediate writeback | Fuse producer/consumer if UB fits; else keep split |
| Frequent host grid chunking | Launch / grid-product overhead | Prefer 1D core-grid (num_aicore Cube / num_vectorcore Vector) + flat task_id |
Low aiv_vec_ratio (~0.75) while MemoryUB is not saturated; larger tiles UB-overflow | Dual DMA paths live in UB | Runtime block_ptr vs masked load: host-split with tl.constexpr so each launch DCE's the other (cases.md § causal_conv1d) |
Colloquial “CUDA utilization” → read Cube/MAC (aic_mac_ratio). Host UB model complements the profiler — see reference.md.
Prioritize fixes by Duration share in kernel_details / op_statistic (largest hotspot first). Low pipe ratios on a dominant kernel usually mean room remains on that pipe.
Change only levers that match the bottleneck; one hypothesis per round. Before tuning, classify the issue: compile failure / UB overflow / grid limit / numeric error / real performance bottleneck — do not treat all five the same way.
peak ≈ memory_multiplier * tiled_elements * dtype_size; comment where the multiplier comes from.fla.utils.ascend_ub_manager (compute_row_tile_block_size, etc.); do not hard-code capacity; keep ~0.75–0.85 safety margin.mem_mult; near 100% and still slow → look at pipe/bandwidth.tl.make_block_ptr + boundary_check; @input_guard for layout — do not emulate arbitrary strides in-kernel.g along T (critical on Ascend): if g is [B, T, HV], host g.transpose(1, 2).contiguous() and load via G_T_CONTIG + stride-1 g_ptr (see g-contiguous-loading.md). Stride-HV gathers in bwd hot loops can be 10×–35× slower than contiguous loads; HV==1 needs no transpose. Match fwd pointer math; keep T_seq before varlen overwrites T.HV==1, layout flags) get separate paths — no expensive hot-loop branches.ASCEND_MAX_GRID_DIM=65535: host-split with iter_axis_launch_chunks, pass *_OFFSET; after varlen slicing, zero the matching offset — never slice and also add a global offset. UB and grid are independent constraints.task_num and schedule with for task_id in tl.range(core_id, task_num, num_core) (or range(pid, total_tasks, num_programs)). Decode task_id → tile indices inside the kernel. One launch, no ASCEND_MAX_GRID_DIM host loop, better load balance when task_num is irregular. Keep do_not_specialize on T / task_num / num_core / dynamic extents.grid=(num_aicore,) via get_device_properties()["num_aicore"]. Vector-bound (conv, layernorm, rotary) → get_multiprocessor_count (num_vectorcore on NPU; A2 is 48 vector vs 24 Cube). Launching a Vector kernel on num_aicore leaves half the vector cores idle.q_ptr = q + …); do not accumulate with in-place ptr += across tasks — Ascend Triton can mis-compile that pattern.i_t, i_b, NT = cdiv(T, BT)) are runtime int32 or narrower. do_not_specialize on T makes NT runtime, but i_t * stride also wraps when T is specialized (packed conv: i_t * BT then offset * D). (NT - 1) * DH_CS wraps past 2³¹ before a trailing .to(tl.int64). Example: DH_CS=HV*K*V, K=V=128, HV=64, BT=64 → overflow at NT>2048 (T>131K). Packed offset * D: T>2³¹/D (D=4096 → T>524K). Cast the index first with tl.cast (not .to on specialized ints): tl.cast(i_t, tl.int64) * BT, tl.cast(i_b, tl.int64) * T, tl.cast(B, tl.int64) * T, tl.cast(NT - 1, tl.int64) * DH_CS. Kernel args B/T are constexpr — B.to(tl.int64) is AttributeError("'constexpr' object has no attribute 'to'"); i_t/i_b can fold to constexpr when NT=1. tl.load(...).to(tl.int64) on cu_seqlens is fine. Never (i_b * T).to(tl.int64) or ((NT - 1) * DH_CS).to(tl.int64).make_block_ptr offsets stay int32: Triton rejects int64 offsets/block_shape. Flattened pointer math (bos * D, t0 * D, i_b * stride) uses int64; pass i_t * BT (int32) as the block row offset. Do not feed t0 into make_block_ptr. Case: causal_conv1d.cu_seqlens → int64 for pointer math: host dtype is often torch.long, but tests also pass int32; load as tl.int64 either way. Loading .to(tl.int32) then (bos * HV + i_hv) * V overflows well before bos hits 2³¹ (HV=32, V=4096 → safe bos ≈ 16K). Pattern: bos, eos = tl.load(cu_seqlens + i_n).to(tl.int64), tl.load(cu_seqlens + i_n + 1).to(tl.int64); T_cur = (eos - bos).to(tl.int32). Non-varlen: bos = tl.cast(i_b, tl.int64) * T (CUDA/repo often writes (i_b * T).to(tl.int64), which still wraps if i_b * T exceeds 2³¹). Alternative when T_cur only needs int32: load bos as int32 but cast the index before the large stride — tl.cast(bos, tl.int64) * HV + i_h then * K. (bos * HV + i_h).to(tl.int64) * K only fixes * K/* V (HV is small); bos * HV itself can still wrap.input_precision='ieee' / allow_tf32=False; mask before exp on gated paths; keep a consistent exp/exp2 base.tl.dot clobbers the left operand: on NPU, tl.dot(lhs, rhs, …) may overwrite lhs in UB (CUDA Triton does not). Any later read of that tile (second lhs, rhs, store) sees corrupted data unless you reload from GM or copy with tile + 0.0 before the first lhs dot. Full per-kernel catalog: cases.md § tl.dot lhs clobber. Symptom: silent numeric drift vs Torch oracle with no compile error.rg 'tl\.dot\(' fla/ops/**/triton_ascend/** — only 8 op files use tl.dot; (2) for each lhs tile, flag lhs→lhs, lhs→rhs/store, or post-dot copy; (3) prefer GM reload for one reuse between stages, + 0.0 for tight multi-dot sequences; (4) re-run tests/ops/test_gdn_kernels.py + op-specific kernel tests.exp2(gs)[:, None] / exp2(gc)[None, :] instead of exp2(gs[:, None] - gc[None, :]) to replace a matrix of exponentials with two vectors. Verify numerics on the target compiler; multiplying by exp2(-gc) can produce materially different Ascend results.if is_tail_chunk that chooses make_block_ptr vs masked tl.load keeps both paths live in UB. Peak UB ≈ sum of both; Vector cannot saturate even when MemoryUB bandwidth is free; larger tiles then fail compile. Host-split the last tile into a second launch with tl.constexpr TAIL_MODE (0 = never tail / block_ptr only, 1 = always masked, 2 = runtime for varlen / NT==1) so each compile DCE's the unused path. Case: causal_conv1d.make_block_ptr whose block end overshoots packed B*T rows faults MTE (DDR address out of range). Use masked load/store on the last chunk, or the constexpr split above so bulk never overshoots. Halo windows (BT+W-1) overshoot even sooner — count the halo in the tail predicate.if USE_INITIAL_STATE or i_t*BT < W still lowers the else and compiles initial_state + … when the pointer is None. Nest: if not FLAG: … elif runtime: … else: ….tl.extract_slice / tl.insert_slice: sliding-window taps without extra GM loads (causal conv). Some triton-ascend versions expose them only via triton.language.extra.cann.extension — shim onto tl if missing. Preloading every tap tile overflows UB; load inside the static_range or one BT+W-1 window + slice.[D, W] → host transpose(0,1).contiguous() to [W, D] for stride-1 channel block_ptr (same idea as G_T_CONTIG). Odd D that cannot be tiled with a power-of-two BD that divides D and BD>=16 falls back to the legacy multi-axis path.o += q@h then intra o += A@v with ACCUMULATE_OUTPUT): if both need the same q (and live set fits), fuse into one kernel — keep b_o / b_A in UB, single store. Profiler cue: two kernels own the op and MTE is high from the intermediate o writeback. If fused peak UB overflows, keep the split; do not force fusion.BT×BT + BT×BV), fix the Cube-aligned outer tile (BV) and autotune the K-slab (BK) rather than host-hardcoding both.tl.debug_barrier only for same-program deps.do_not_specialize=['T'] (and other dynamic launch extents); kill runtime branches with tl.constexpr / triton.heuristics.num_warps / num_stages anywhere: NPU does not support them. Omit from kernel call sites, autotune config dicts, and wrappers. Do not leave them commented-out “for CUDA parity”; delete them.prepare_chunk_indices / prepare_chunk_offsets. With 1D core-grid, flatten over total_chunks and map global_t → (i_n, i_t) via chunk_offsets (largest i_n with chunk_offsets[i_n] <= global_t). Tests cover empty tails, non-aligned lengths, multi-length, and fixed/varlen equivalence.Failure modes and repo paths: reference.md. Detailed past cases: cases.md.
Each round, in order:
triton_ascend.aic_metrics and confirm Duration/pipe/UB move as expected.Prefer: tests/ops/test_gdn_kernels.py, tests/ops/test_solve_tril.py, tests/modules/test_conv.py (causal_conv1d), tests/utils/test_ascend_ub_manager.py, python -m benchmarks.ops.verify --op <op> --base <ref> (--gate-k is a quick signal only).
After re-profile, report:
fla-optimization-loop stop criteria)Generalizable fixes discovered during optimization belong in this skill (SKILL.md, references/reference.md, or references/cases.md) in a separate doc commit — not bundled into a perf PR.
num_warps / num_stages in Ascend kernel launches, autotune configs, or wrappersnum_aicore Cube / num_vectorcore Vector); host-split offsets not double-counted with varlen; task-loop pointers rebound each iterationblock_ptr vs masked DMA: constexpr-split so bulk DCE's the unused path; tail DMA does not overshoot packed B*T (include halo)or-ed with runtime checks (None ptr must not compile)g uses G_T_CONTIG when [B,T,HV] (see g-contiguous-loading.md); tail boundary_checkACCUMULATE_OUTPUT writeback when UB allows)tl.dot left-hand tiles: GM reload or tile + 0.0 before first lhs dot (post-dot copy invalid); see cases.md § tl.dot catalogNT, i_t, i_b, B, program IDs) via tl.cast(..., tl.int64) before stride / BT / D multiply — including packed offset * D. Not gated on do_not_specialize. Do not call .to(tl.int64) on specialized kernel args (constexpr has no .to)bos/eos from cu_seqlens loaded as tl.int64; T_cur = (eos - bos).to(tl.int32) only; non-varlen tl.cast(i_b, tl.int64) * Tmake_block_ptr offsets/block_shape stay int32 (i_t * BT); int64 is only for flattened ptr + offset * strideuse_g True/False with g=None reference) when PR touches gated and ungated pathstest_*.pynum_warps / num_stages on Ascend paths (unsupported; not a tuning lever)is_tail_chunk (or similar) between block_ptr and masked DMA — both stay in UBnum_aicore (half the vector cores idle on A2)if CONSTEXPR_FLAG or runtime: around an optional pointer — else still compiles when the ptr is NoneB.to(tl.int64) / i_t.to(tl.int64) on specialized or folded constexpr ints (constexpr has no .to); use tl.castt0 as make_block_ptr offsets (offsets/block_shape must be int32)scripts/profile_npu.py, scripts/analyze_profile.pyg stride-1 loading (G_T_CONTIG): g-contiguous-loading.mdconstexpr .to, int64 block_ptr offsets): TRAPS.mdnpu_prof/ (new collection must use the generic scripts)© fla-org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (scripts, references) in .agents/skills/fla-ascend-performance of fla-org/flash-linear-attention.
Open the folder on GitHubat commit b8ff848
Fla Ascend Performance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Fla Ascend Performance this skillfla-org/flash-linear-attention | 5.8k | — | ~5.6k | Automated safety check: Pass | MIT | |
| JS Perf InvestigationSAP/project-foxhound | 180 | 1 repos | ~4.1k | Automated safety check: Pass | GPL-3.0 | |
| Data Profileraspi6246/Claude-Code-Skills-for-Academics | 157 | — | ~2k | Automated safety check: Pass | None | |
| Data Scientistmagnus919/agent-skills | 111 | — | ~4.1k | Automated safety check: Pass | MIT | |
| Constant Time Testingtrailofbits/skills | 7.4k | — | ~5.2k | Automated safety check: Pass | CC-BY-SA-4.0 | |
| Analyzing Network Flow Data With Netflowmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~543 | Automated safety check: Pass | Apache-2.0 |
SAP/project-foxhound
Structured performance opportunity investigation for SpiderMonkey (the Firefox JavaScript engine).
aspi6246/Claude-Code-Skills-for-Academics
Systematic dataset profiling protocol for empirical research.
magnus919/agent-skills
A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…
trailofbits/skills
Measures timing side channels in cryptographic implementations by running them, using dudect for statistical analysis and Timecop over Valgrind for dynamic tracing.
mukul975/Anthropic-Cybersecurity-Skills
Parse NetFlow v9 and IPFIX records to detect volumetric anomalies, port scanning, data exfiltration, and C2 beaconing patterns.
GPTomics/bioSkills
Turns a shotgun profiler table (MetaPhlAn relative abundance, Bracken counts, HUMAnN function tables) into honest figures and defensible community statistics with phyloseq, vegan, microViz, and…
fla-org/flash-linear-attention
Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness.
fla-org/flash-linear-attention
Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…
fla-org/flash-linear-attention
Guidelines for kernel correctness testing and coverage in fla/ops/ and related modules, including common Triton grid/addressing pitfalls.
fla-org/flash-linear-attention
Contract-first design and coverage discipline for FLA kernel and numerical changes.
fla-org/flash-linear-attention
Workflow for FLA backend dispatch decorators and backend implementations.
fla-org/flash-linear-attention
FLA KDA kernel workflow and public technical notes. An agent skill from fla-org/flash-linear-attention.
Categories
Guidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo. Fla Ascend Performance is an agent skill from fla-org/flash-linear-attention. Guidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo.
Fla Ascend Performance fits situations like: working on NPU profiling; kerneldetails/opstatistic; fla tritonascend backends (ops; G transpose stride-1.
Run `npx skills add fla-org/flash-linear-attention --skill fla-ascend-performance -a claude-code`. Or copy the skill folder (.agents/skills/fla-ascend-performance in fla-org/flash-linear-attention) into .claude/skills/fla-ascend-performance in your project. Claude Code loads it when a task matches its description.
Run `npx skills add fla-org/flash-linear-attention --skill fla-ascend-performance -a codex`. Or copy the skill folder (.agents/skills/fla-ascend-performance in fla-org/flash-linear-attention) into .agents/skills/fla-ascend-performance in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fla-org/flash-linear-attention --skill fla-ascend-performance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fla-ascend-performance, .gemini/skills/fla-ascend-performance, .github/skills/fla-ascend-performance and .opencode/skills/fla-ascend-performance in your project.
Going by SKILL.md and its folder, Fla Ascend Performance needs Python for the scripts in its folder and the command-line tools its instructions call (python and rg). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Fla Ascend Performance is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Fla Ascend Performance: JS Perf Investigation (SAP/project-foxhound, 180 stars), Data Profiler (aspi6246/Claude-Code-Skills-for-Academics, 157 stars), Data Scientist (magnus919/agent-skills, 111 stars) and Constant Time Testing (trailofbits/skills, 7.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
fla-org (a GitHub organization) maintains it in fla-org/flash-linear-attention, which has 5,828 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 6, 2026.
Source: fla-org/flash-linear-attention on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.