Agent skill

Lora Routing

by sorryhyun in sorryhyun/anima_lora

The LoRA-family three-axis routing surface (usemoestyle / routeperlayer / routersource) — the variant matrix, per-variant module details (LoRA/T-LoRA/Hydra), and GlobalRouter mechanics.

MITAuto-check passedAI & LLM Engineering

Install Lora Routing

skills CLI
$ npx skills add sorryhyun/anima_lora --skill lora-routing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sorryhyun/anima_lora lora-routing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sorryhyun/anima_lora.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/lora-routing .claude/skills/lora-routing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
lora-routing
GitHub stars
125
Token cost
~1.3k tokens
SKILL.md length
499 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

The LoRA-family three-axis routing surface (usemoestyle / routeperlayer / routersource) — the variant matrix, per-variant module details (LoRA/T-LoRA/Hydra), and GlobalRouter mechanics.

  • Tasks that involve Fine-tuning
  • SKILL.md covers Variant matrix, LoRA variants and GlobalRouter (network-level…
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Lora Routing is an agent skill from sorryhyun/anima_lora. The LoRA-family three-axis routing surface (usemoestyle / routeperlayer / routersource) — the variant matrix, per-variant module details (LoRA/T-LoRA/Hydra), and GlobalRouter mechanics. Load before adding/changing a LoRA variant, touching routing code in networks/loraanima/ or loramodules/, or debugging router behavior.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Fine-tuning. The repository describes itself as: optimized anima lora training script. The licence is MIT.

When your agent uses it

  • Tasks that involve Fine-tuning

Example prompts

  • “/lora-routing”

What it can do on your machine

Read from SKILL.md and the folder at commit d16b651. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Lora Routing loads about 1.3k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 499 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sorryhyun/anima_lora at commit d16b651, republished under its MIT licence (© sorryhyun). 499 words, ~1,301 tokens.

Download SKILL.mdSave it as .claude/skills/lora-routing/SKILL.md (or your agent's skills folder).
name
lora-routing
description
The LoRA-family three-axis routing surface (use_moe_style / route_per_layer / router_source) — the variant matrix, per-variant module details (LoRA/T-LoRA/Hydra), and GlobalRouter mechanics. Load before adding/changing a LoRA variant, touching routing code in networks/lora_anima/ or lora_modules/, or debugging router behavior.

LoRA-family routing: three-axis surface, variants, GlobalRouter

Variant matrix

The three axes (use_moe_style / route_per_layer / router_source) are consumed by lora_anima/config.py::LoRANetworkCfg.from_kwargs and dispatched by networks/__init__.py::resolve_network_spec. Legal values:

AxisValues
use_moe_styleFalse / "shared_A"
route_per_layertrue / false
router_source"none" / "input" / "sigma" / "fei" / "crossattn_emb" (empty / unset → "none")

Two constraints, both raising in from_kwargs: "input" requires route_per_layer=True (no per-Linear input signal reaches a network-level router); "crossattn_emb" requires route_per_layer=False.

Variants that exist as cells in this matrix:

Variantuse_moe_styleroute_per_layerrouter_sourceNetwork module / path
Plain LoRA / T-LoRAFalse—"none"lora_anima + lora_modules/lora.py
HydraLoRA (paper)"shared_A"True"input"lora_anima + lora_modules/hydra.py
σ-router on Hydra"shared_A"True"sigma"same
FEI-on-Hydra"shared_A"True"fei"same
Text-routed Hydra"shared_A"False"crossattn_emb"lora_anima + GlobalRouter (pools + LN on the cross-attn text vector)

The "crossattn_emb" cell routes the whole pool by prompt content (pooled post-LLM-adapter text features) instead of σ/noise-frequency: the network-level GlobalRouter reads the same vector the DiT cross-attends to, fired per cond/uncond branch via set_crossattn_routing (train, train.py) / set_hydra_crossattn (inference, library/inference/generation.py), broadcasting to the standard _routing_weights slot.

Pre-plan2 metadata stamps (ss_use_hydra, ss_use_fei_router) no longer load; the stamps are now ss_use_moe_style / ss_route_per_layer / ss_router_source. use_moe_style="independent_A" (the former stacked-experts / FeRA layout) is no longer a legal value — passing it raises a ValueError (networks/lora_anima/config.py::_as_moe_style).

LoRA variants

All live in networks/lora_modules/. Stack freely via toggle flags in configs/methods/lora.toml.

  • LoRA (lora.py::LoRAModule) — Classic low-rank: y = x + (x @ down @ up) * scale * multiplier.
  • T-LoRA — Not a separate class. A _timestep_mask buffer on LoRAModule (registered in base.py) is rebound to a shared live-updated mask by lora_anima/network.py::LoRANetwork.set_timestep_mask. Effective rank varies with denoising step via a power-law schedule. Training-only — inference runs full rank at every t (baking into DiT is bit-equivalent). See docs/methods/timestep_mask.md.
  • HydraLoRA (hydra.py) — MoE-style multi-head routing: shared lora_down + per-expert lora_up_i heads, layer-local router on the adapted Linear's input (router_source="input") or σ-features / FEI features ("sigma" / "fei"). With route_per_layer=False the per-layer router drops out for a network-level GlobalRouter fed σ-features, FEI, or pooled cross-attn text (router_source="crossattn_emb"). Requires cache_llm_adapter_outputs=true. Produces a *_moe.safetensors sibling for router-live inference. See docs/methods/hydra-lora.md.

ReFT was removed from the live tree on 2026-06-08 and downgraded to a bench probe — module, configs, docs and a re-integration map live in bench/reft/ (INTEGRATION.md + impl/).

Show full SKILL.md (153 more words)Show less

GlobalRouter (network-level routing)

lora_anima/routers.py::GlobalRouter (re-exported from network.py for back-compat) — Linear(F_in → H) → ReLU → Linear(H → E) → softmax/τ. Built when cfg.route_per_layer=False and cfg.use_moe_style != False. Final layer is zero-init so step-0 gates are uniform; warmup is the symmetry-breaker. Under router_source="crossattn_emb" the router is built with apply_layer_norm=True and input_dim=CROSSATTN_EMB_DIM; its forward RMS-pools a raw (B, L, D) text tensor over the sequence axis and LayerNorms (parameterless) before the MLP — no extra state_dict keys, on/off is deterministic from router_source.

Hook site: LoRANetwork.set_fei(z_t) runs the FEI computation (via library/runtime/fei.py) and the router once, then writes the resulting (B, num_experts) tensor by reference into each routing-aware module's _routing_weights buffer. One Python-level write propagates to every adapted Linear that step — hence the failure mode to watch for: router collapse takes every layer down together.

Training-loop call: train.py fires network.set_fei(noisy_model_input) at the per-step σ/FEI hook block when the cfg has route_per_layer=False and router_source="fei". Inference: library/inference/generation.py mirrors the same call before each Euler step.

© sorryhyun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/lora-routing of sorryhyun/anima_lora.

Open the folder on GitHubat commit d16b651

Compare with similar skills

Lora Routing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Lora Routing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Lora Routing this skillsorryhyun/anima_lora125—~1.3kAutomated safety check: PassMIT
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0
Train RlOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0
Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide1.7k—~830Automated safety check: PassApache-2.0
Dataset Evaluationawslabs/agent-plugins9161 repos~1.3kAutomated safety check: PassApache-2.0
Train SftOpenPipe/ART11k—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Qwopus27b Rl Training

    R6410418/Jackrong-llm-finetuning-guide

    Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

    1.7k GitHub stars~830 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    916 GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Sft

    OpenPipe/ART

    SFT training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Fine Tuning With Trl

    Orchestra-Research/AI-Research-SKILLs

    Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.

    13k GitHub starsUsed in 6 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed

More from sorryhyun/anima_lora

All 12 skills in this repo
  • Model Catalog

    sorryhyun/anima_lora

    The model catalog (library/downloads.py) — one Asset row per weight (repo, files, destination, installed probe), packs, resolve() name order, and the rule that loaders import their default paths…

    125 GitHub stars~502 tokensUpdated today
    Auto-check passed
  • Qwen21

    sorryhyun/anima_lora

    Qwen-Image-2.1 LoRA line (NOT Anima) — running cache/train through the daemon, make gui-qwen, the CacheRequest/TrainRequest flag surface and how to add a field, model-dir resolution, cache layout…

    125 GitHub stars~1.9k tokensUpdated today
    Auto-check: notes
  • Anime Tools

    sorryhyun/anima_lora

    The trainer ↔ animetools boundary — what the curation split moved out, the typed request/stage API the make targets build, the git-pin dev loop and its stale-venv trap, and the tests that guard the…

    125 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Bucketing

    sorryhyun/anima_lora

    Free-fit native-shape bucketing — the token bands per edge tier, tier choice at preprocess time, the compiledynamicseq coupling and per-tier graph budget, and why training never needs --targetres.

    125 GitHub stars~932 tokensUpdated today
    Auto-check passed
  • Captions

    sorryhyun/anima_lora

    Caption pipeline — position-clause grammar (never hand-split a caption), make caption-autotag modes, make caption-position (v2 rewrite rules and gates), and the preprocess-stage wiring for both.

    125 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Custom Nodes

    sorryhyun/anima_lora

    The ComfyUI node map — which node lives in which standalone repo vs in-tree under customnodes/, where each is symlinked, and the vendor-sync rule for the vendor/ subsets.

    125 GitHub stars~470 tokensUpdated today
    Auto-check passed

Questions about Lora Routing

What does Lora Routing do?

The LoRA-family three-axis routing surface (usemoestyle / routeperlayer / routersource) — the variant matrix, per-variant module details (LoRA/T-LoRA/Hydra), and GlobalRouter mechanics. Lora Routing is an agent skill from sorryhyun/anima_lora. The LoRA-family three-axis routing surface (usemoestyle / routeperlayer / routersource) — the variant matrix, per-variant module details (LoRA/T-LoRA/Hydra), and GlobalRouter mechanics.

When should I use Lora Routing?

Lora Routing fits situations like: tasks that involve Fine-tuning.

How do I install Lora Routing in Claude Code?

Run `npx skills add sorryhyun/anima_lora --skill lora-routing -a claude-code`. Or copy the skill folder (.claude/skills/lora-routing in sorryhyun/anima_lora) into .claude/skills/lora-routing in your project. Claude Code loads it when a task matches its description.

How do I install Lora Routing in Codex?

Run `npx skills add sorryhyun/anima_lora --skill lora-routing -a codex`. Or copy the skill folder (.claude/skills/lora-routing in sorryhyun/anima_lora) into .agents/skills/lora-routing in your project. Codex loads it when a task matches its description.

Can I use Lora Routing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sorryhyun/anima_lora --skill lora-routing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lora-routing, .gemini/skills/lora-routing, .github/skills/lora-routing and .opencode/skills/lora-routing in your project.

What does Lora Routing need to run?

SKILL.md names no scripts, command-line tools or credentials: Lora Routing is instructions for the agent only.

Does Lora Routing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Lora Routing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Lora Routing use?

Lora Routing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Lora Routing use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Lora Routing?

Skills that share tags, products or a category with Lora Routing: Sentence-Transformers Training Router (huggingface/skills, 11k stars), Train Rl (OpenPipe/ART, 11k stars), Qwopus27b Rl Training (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars) and Dataset Evaluation (awslabs/agent-plugins, 916 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Lora Routing?

sorryhyun (a GitHub user) maintains it in sorryhyun/anima_lora, which has 125 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 10, 2026.

Source: sorryhyun/anima_lora on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.