Research workflow for distilling ML/AI papers into modelless inference primitives, freeze/thaw runtime patterns, latent-space operations, AND model-based training plans across the multi-repo stack.
Install the "research" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/research into .claude/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add katopz/katgpt-rs --skill research -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
GitHub CLI
$ gh skill install katopz/katgpt-rs research --agent codex
Project scope by default (.agents/skills/); add --scope user for a personal install.
Install the "research" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/research into .agents/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add katopz/katgpt-rs --skill research -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
GitHub CLI
$ gh skill install katopz/katgpt-rs research --agent cursor
Project scope by default (.agents/skills/); add --scope user for a personal install.
Install the "research" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/research into .cursor/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add katopz/katgpt-rs --skill research -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
GitHub CLI
$ gh skill install katopz/katgpt-rs research --agent gemini-cli
Project scope by default (.agents/skills/); add --scope user for a personal install.
Install the "research" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/research into .gemini/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
GitHub CLI
$ gh skill install katopz/katgpt-rs research
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add katopz/katgpt-rs --skill research -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "research" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/research into .github/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add katopz/katgpt-rs --skill research -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
GitHub CLI
$ gh skill install katopz/katgpt-rs research --agent opencode
Project scope by default (.agents/skills/); add --scope user for a personal install.
Install the "research" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/research into .opencode/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
research
GitHub stars
135
Token cost
~20k tokens
SKILL.md length
9,455 words
Files
3
Skills in repo
8
Repo updated
First seen
Licence
MIT
At a glance
Research workflow for distilling ML/AI papers into modelless inference primitives, freeze/thaw runtime patterns, latent-space operations, AND model-based training plans across the multi-repo stack.
Works in 6 steps: Read & classify → Distill fundamentally — fuse, don't… → If gain/GOAT, plan it → …
Reading arxiv papers
SKILL.md covers When to use, Repos, Commercial strategy (inline… and Pre-flight — MANDATORY before…, plus 6 more sections
Calls git; reaches r.jina.ai and github.com
What it does
Research is an agent skill from katopz/katgpt-rs. Research workflow for distilling ML/AI papers into modelless inference primitives, freeze/thaw runtime patterns, latent-space operations, AND model-based training plans across the multi-repo stack. Use when reading arxiv papers, deciding which repo a paper belongs in, creating .research/ notes or .plans/ files, implementing modelless inference primitives, or routing training-vs-inference insights. Enforces the commercial strategy (public engine / private runtime / private chain / private neuron-db / private…
Its SKILL.md is about 20k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `templates.md` and `vocab.md`).
It sits in AI & LLM Engineering, covering Fine-tuning. It works with arXiv. The repository describes itself as: A neuro-symbolic micro-Transformer with speculative decoding, constraint pruning, recurrent attention, and adaptive test-time scaling — built in Rust. The licence is MIT.
When your agent uses it
Reading arxiv papers
Deciding which repo a paper belongs in
Creating .research/ notes
Implementing modelless inference primitives
Example prompts
“/research”
Workflow steps
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 7e47c8c. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
Shell commands in SKILL.md call:
git
From the folder's file list and the shell code blocks in SKILL.md.
Network
Hosts in commands or code, which the agent is likely to contact:
r.jina.ai
github.com
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
Research loads about 20k tokens when it runs. Until then it costs about 242 tokens; SKILL.md has 9,455 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~242
When it runs· the whole SKILL.md, loaded when a task matches
~20k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Download SKILL.mdSave it as .claude/skills/research/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
research
description
Research workflow for distilling ML/AI papers into modelless inference primitives, freeze/thaw runtime patterns, latent-space operations, AND model-based training plans across the multi-repo stack. Use when reading arxiv papers, deciding which repo a paper belongs in, creating .research/ notes or .plans/ files, implementing modelless inference primitives, or routing training-vs-inference insights. Enforces the commercial strategy (public engine / private runtime / private chain / private neuron-db / private training / private SDK facade / private product-domain / public inference-substrate riir-infer / public decision-serving riir-reflex), three-track system (modelless inference + self-adaptive runtime + model-based trained weights — all three exist across the stack, not just riir-train), latent-to-latent preference, and freeze/thaw-over-fine-tuning rule. Every task ends with a mandatory Claude verdict ping-pong (§5) before any file is committed.
Research Workflow — Modelless Inference, Freeze/Thaw, Latent-to-Latent
This repo (katgpt-rs) + riir-ai (freeze/thaw runtime + adaptive NPCs + game systems) + riir-chain (neuro-symbolic chain transport, LatCal) + riir-neuron-db (NeuronShard, BLAKE3/Merkle, freeze/thaw envelope) ship runtime + latent-space operations. The inference substrate (weights, quant formats, kernels, encoders, loaders) lives in riir-infer; decision-engine serving + comparison (calibration, abstention, arenas) lives in riir-reflex — both opened public 2026-09-23 as the first sanctioned Research-003 exceptions. Training-method research lives in riir-train. If a paper's value is its training loop → riir-train (see §3.5 Path 0.5 — applicable training papers get a Plan, not a lazy redirect). If its value is a latent-space insight, a routing trick, a freeze/thaw pattern, a chain-commitment bridge, a neuron-shard primitive, or a modelless inference primitive → distill here. If it's about how trained weights are REPRESENTED and run (quant, kernels, encoder forwards, loaders) → riir-infer; if it's about how a MODELLESS decision is SERVED (calibration, abstention, comparison arenas) → riir-reflex; if it's a TRAINED-ENCODER serving arm (encoder lane product, per-suite winners/vessels, serving posture) → riir-rethink (private forever — reflex gets at most an owner-gated public comparison lane).
When to use
Reading/fetching/summarizing ML/AI/systems papers · deciding which repo a paper belongs to · creating .research/ notes or .plans/ files · implementing modelless inference primitives · designing freeze/thaw cycles, adapter hot-swap, runtime adapter routing · designing latent-to-latent ops (dot-product projections, sigmoid gating, manifold geometry, spectral methods) · designing MMORPG-scale game AI (thousands of concurrent NPCs, 20Hz tick, fog-of-war, emergent social/economic behavior).
Do NOT activate for: pure refactor, bug fixes with no research angle, or ordinary feature work.
Repos
Canonical count + membership: katgpt-rs/AGENTS.md §"Repo count" — the
product/distillation set is 7 (first block below); the workspace is 27
contract repos at the 2026-10-07 read (all with a root BOUNDARY.md; the
Instinct/Rethink carve added two 2026-10-03 — this file's typed counts have
rotted before, so derive the set, never trust this or any number here):
bash
cd /Users/katopz/git && for d in */; do
case "$d" in *.wprobe/) continue;; esac # probe artifacts carry BOUNDARY.md + .git but are not repos (naive derive returns 28)
[ -f "$d/BOUNDARY.md" ] && [ -d "$d/.git" ] && echo "${d%/}"
done
The product/distillation set (7):
katgpt-rs/ — public MIT engine. Generic modelless inference primitives. No game/chain/shard IP.
riir-ai/ — private game product. Freeze/thaw runtime, self-learn, game systems. Hosts the .docs/ moat book.
riir-train/ — private training vault. As of 2026-08-06: actively pursued, not lazily redirected. Applicable training papers get a Plan in riir-train/.plans/ per §3.5 Path 0.5. read_file riir-train/.docs/02_pipelines/training_data_pipeline.md before any training-paper verdict.
riir-dapps/ — private dApp layer (added 2026-08-20). Game outcome → generic chain settlement (Settlement, MultiClaimEscrow composition) + the KAT ledger/service SERVER plane (kat_ledger, epoch settle, payment rails, CF worker/DO shell). Route settlement-composition + service-economy papers here, NOT riir-chain.
**The wider workspace (the other 20 — route here only when the insight is
inseparable from that surface; the first three entries are first-class PUBLIC
distillation targets, not last resorts):
riir-refine/ — private multi-domain code healer. Fusion priority #2 (see the ladder below) — the workspace's densest latent-state consumer per LOC; healer-domain distills file their .research/ notes HERE (renamed riir-refine -> riir-refine 2026-10-01) (MOAT row below).
riir-infer/ — PUBLIC LLM-inference substrate (carved from riir-ai 2026-09-22; opened public 2026-09-23 — the first sanctioned Research-003 exception). Workspace: riir-infer-core at the repo root (CPU substrate: weights/quant/architectures/loaders — quant/ EXL3-trellis lane, deltanet/ GDN lane, Bonsai2 rotation, gguf_loader.rs/safetensors_loader.rs/tokenizer.rs, simd/) + crates/riir-infer-gpu (wgpu/CubeCL GPU layer) + crates/riir-infer-laya (the pinned-checkpoint encoder lane: CPU + macOS Metal + non-macOS CUDA backends — reflex's comparison-lane substrate). Zero riir- deps — upstream of the engine.* Quant/weight-format/kernel/encoder/loader papers distill HERE (a real target; its .research/ is created on first use).
riir-reflex/ — PUBLIC decision-engine serving + comparison + contribution (opened public 2026-09-23, same day). Serves decision_wire (choice/score/noul, abstention first-class) modellessly: hashed-feature embedding (embed.rs) → pick_domain corpus routing → Lz4FlexDrafter corpus-is-the-model option scoring → sigmoid normalization → SigmoidGateCalibrator → fused abstain (CorpusDistanceGate). GOAT gates G1–G5 (G5 laya-parity must pass before ANY published laya number); the honest-metrics harness (harness/, 9 dataset suites, corpus-cap levers) + comparison lanes (lanes/: laya CPU/Metal/CUDA, CLM, GLiNER, AgentJev). Calibrated-decision-serving / abstention / selective-prediction / comparison-arena papers route here. The laya lane SUBSTRATE lives in riir-infer — reflex consumes + measures it.
riir-instinct/ — PUBLIC open lane (carved from the moat 2026-10-03, Instinct/Rethink split): the L1→L3 decision-serving stack crate — suites, the lane-backend seam (riir_instinct::server::{LaneBackend, ExtBoots, install_ext_boots} + the AnySuiteServer::Ext seat), arena game heads, and the one-definition home of the escalation grammar (riir_instinct::arsenal::EscalateSpec; riir-refine Plan 202 R1 — rethink byte-pins it, never forks it). Open-stack/lane-seating/escalation-grammar/arena-head papers route here; the enforcement + product posture live rethink-side.
riir-rethink/ — PRIVATE FOREVER (the carve's moat, seeded from riir-instinct/moat/): the trained-encoder decision product — L3 encoder serving rung (serving-weights sizing: F16/Q8 adopted/Q4 FAIL), ESC enforcement (the manifest escalate table is the only ARMING surface; RIIR_RETHINK_ESC=0 demote-only; rolling-rate guard), the HOSTED-ONLY vessel reader (the vessel feature never joins a public build — the A10 moat law), the decstat economy-capture plane, and the ESC-depth research line (.research/001 population-scaling). ESC/escalation-lane/depth-lane/population-scaling/serving-economy papers route here. Visibility: git@github.com:gist-rs/* private; nothing here may leak through a public-repo distill.
riir-auth/ — private. Latent-space session recording (SessionFingerprint → NeuronShard) + 3-pillar ban quorum + the account_key Ed25519 identity substrate. Behavioral-identity/fingerprinting/escalation papers route here.
riir-esp32/ — private POC/fun, not prod: the ESP32 Satellite device tier (esp-hal firmware, on-device crypto timing). Device-tier papers only.
riir-viewbridge/ — private Unity FFI seam, PARKED 2026-09-03 (Bevy is the shipping render path). FFI raw/latent-wall papers only; expect no consumer until unfreeze.
katgpt-web/ — PUBLIC explainer site ("The Anatomy of KatGPT-RS"). Presentation surface only — nothing private may ever appear here; never a distillation target.
riir-shader/ — private shader surface (workspace member per scripts/repo_set.txt; no established research routing yet).
(riir-armageddon/ sat in the product set until it was retired 2026-09-02, owner
act — the directory is gone; do not route to it.)
Also on disk, never routed by name: riir-mmorpg-examples, riir-llm,
mmorpg-editor, mmorpg-remake, mmorpg-remaster, riir-reflexer,
reflex-site — consumers/leaf/presentation surfaces, targets of last resort
(riir-mmorpg-examples is the reference game consumer; reflex-site is the
public bench-publish front).
Routing rule of thumb: if it's about how a shard is structured/committed/frozen/consolidated/retrieved/projected → riir-neuron-db. If it's about how a shard crosses quorum or bridges to LatCal fixed-point → riir-chain. If it's about code retrieval / fix trajectories / rule corpora / healer selection → riir-refine. If it's about settlement composition or the service economy → riir-dapps (economy governance/tuning → riir-dao; the client/wire protocol → riir-kat). If it's about how trained weights are represented and run — quant formats, kernels, encoder forwards, loaders → riir-infer. If it's about how a MODELLESS decision is served — calibration, abstention, comparison arenas → riir-reflex; a lane that SERVES product decisions from a trained encoder (encoder lane product, per-suite winners/vessels, serving-weights posture) is riir-rethink's charter even when the proposal says "Reflex" — reflex gets at most an owner-gated public comparison lane. If it's about the escalating decision ladder, suite/arena seating, or the open serving stack's grammar → riir-instinct (open stack + EscalateSpec one-definition) / riir-rethink (private enforcement + depth economics; split by the carve law: grammar and stack in instinct, product posture and enforcement in rethink, lockstep edits in ONE family).
⛔ Family ≠ repo (the 2026-10-07 encoder-lane instance, 6th of the challenge-rescue class; the record lives in the private tier): the ladder reflex→instinct→rethink is a product family spanning THREE repos, and input documents routinely name "the Reflex family/stack" as host. Route by RUNG, never by the document's own name for its host: modelless floor + public arenas = reflex; open grammar/seam/arena heads = instinct; trained-encoder serving + enforcement + economy = rethink. The decisive question for any encoder/decision-lane proposal: is this a public comparison lane or a private serving arm?
What = public. How = private. Training how → riir-train. Runtime how → riir-ai. Chain how → riir-chain. Shard how → riir-neuron-db. When unsure → default private (safe to keep private; never safe to un-leak public). Toy 2D games (bomber, monopoly, go, fft) are NOT product IP — public. *_runtime suffix = private GOAT composition layer; bare-name = public primitive. FV moat: ~79 Lean 4 theorems across 4 .proofs/ instances.
Public exceptions (2026-09-23, owner-sanctioned Research-003 amendments): riir-infer (the model-based inference substrate — weights/quant/kernels/loaders; the substrate tier of track c) and riir-reflex (decision-engine serving + comparison arenas). They are routing targets with their own MOAT rows (§1.6), not leaks — runtime/chain/shard/train IP stays private.
Pre-flight — MANDATORY before any verdict/file
Do all five before creating any file:
read_file the product-set READMEs (katgpt-rs, riir-ai, riir-chain, riir-neuron-db, riir-train, riir-game-sdk, riir-dapps) + riir-ai/.docs/README.md + riir-refine/AGENTS.md (the healer — fusion priority #2) + riir-infer/AGENTS.md + riir-reflex/AGENTS.md (the public inference-substrate + decision-serving surfaces) + riir-instinct/AGENTS.md + riir-rethink/AGENTS.md (the ladder split — reading these two is what separates a reflex public comparison lane from a rethink private serving arm) + (for training papers) riir-train/.docs/02_pipelines/training_data_pipeline.md — defines scope boundaries. Skipping = #1 cause of false Super-GOAT claims + false PASS on training papers.
list_directory every .research/ folder in the workspace — derive it (ls -d /Users/katopz/git/*/.research), never trust a typed count: 15 today (katgpt-rs, riir-ai, riir-auth, riir-chain, riir-refine, riir-dao, riir-infer, riir-instinct, riir-neuron-db, riir-reflex, riir-rethink, riir-shader, riir-train, riir-viewbridge, seal-online-remaster); create missing ones on first use.
list_directory the 6 runtime/chain/db/substrate/serving src trees — module names are vocab. Skipping = #2 cause of false Super-GOATs.
web_search for published prior art on the paper's headline technique (see §4). Skipping = #4 false-novelty failure mode. Internal-first rule (GRAM/058 lesson, 2026-09-25): before any web search, grep every workspace .research/ folder for the paper's arXiv ID AND its title verbatim — a prior distillation is the controlling constraint, and a content grep that stops at page 1 of matches does not find it (Research 058 was missed exactly this way at note-590 round 1; caught only by the verdict reviewer).
grep ALL repos for existing training/model-based/self-adaptive code — the system is three-track, not modelless-only. Skipping = #5 false-PASS mode (canonical failure: arXiv:2511.18538 Code LLM survey was falsely PASSed because the agent only checked riir-train for training code, missing quest_grammar's LoRA training in riir-ai + TernaryDraftModel in riir-refine + self_evolve). Grep patterns: *_training.rs|*_train*.rs|LoRA|SFT|GRPO|DPO|ternary|TernaryDraftModel|self_evolve|with_weights|\.bits across riir-ai/crates/**/src/, riir-refine/src/, riir-neuron-db/src/. Training pipelines ship OUTSIDE riir-train — the three tracks are: (a) modelless inference, (b) self-adaptive runtime latent updates, (c) model-based trained weights.
list_directory the src trees (vocab; skipping causes the vocabulary-miss class of false verdicts — mechanisms ship under non-obvious names):
The highest-value Super-GOATs come from fusing 2–3 papers/primitives into a novel combination, not from direct-mapping a single paper. Always grep .research/ + .plans/ for the 2–3 closest cousins before verdict, and ask: "what does paper × note A × note B produce that none of them alone can?"
The repos are NOT equal-weight fusion surfaces. When a paper fuses into several, evaluate and file in THIS order — the higher-priority surface wins plan/issue filing priority for the same primitive:
Game / MMORPG runtime (riir-ai + riir-game-sdk + product repos) — the #1 investment.
riir-refine healer (code/perf/sec fixing: corpus, retrieval, fix trajectories, score-bench, rustc_errors playbooks, domain router) — the #2 investment and the workspace's densest latent-state consumer per LOC.
Training / fine-tuning / LoRA (riir-train + in-repo pipelines; recipes and adapters — full pretraining is out of scope).
The ladder orders EFFORT, not existence: lower-priority fusions still get recorded (note + issue), but a higher-priority surface carrying the same mechanism gets the plan first. The standing per-verdict check this ladder encodes: "did I ask what this does for the healer?" — canonical failure #5 (below) is what happens when that question is never asked.
Patterns that ship here:
Latent-to-latent operations — dot-product projections, cosine retrieval, sigmoid-gated routing, manifold geometry, spectral methods on activations. Prefer operating on latents over decode→re-encode. Fuse with freeze/thaw to version direction vectors; with self-learn to update them from runtime curiosity.
Freeze/thaw patterns — versioned snapshots, atomic hot-swap, lock-free reads, BLAKE3-checked reload, per-entity personality divergence. Fuse with adapter routing to dispatch by latent similarity; with self-learn to checkpoint emergent personalities.
Runtime adapter routing — selecting between frozen adapters by state/objective/context (Dynamic Pair, Polytope, dMoE — inference-time, zero training). Fuse with freeze/thaw to make the pool versioned + committed; with bandits to learn routing online.
Self-learn / adaptive CoT — runtime curiosity, entropy exploration, collapse detection/recovery, latent prediction SSL, trajectory folding. No backprop. Fuse with MMORPG-scale game AI for thousands of NPCs each with independent curiosity; with freeze/thaw to checkpoint.
MMORPG-scale game AI — thousands of concurrent NPCs, 20Hz tick, fog-of-war, zone attention, emergent social/economic behavior. Latent ops must batch; raw sync must stay bit-identical.
Super-GOAT factory modules — grep FIRST
The highest-value latent Super-GOATs cluster in twelve module trees. list_directory these explicitly even if the paper looks pure-training — vocab mismatch is the #3 cause of false verdicts:
Module
What ships
Super-GOAT angle
katgpt-rs/crates/katgpt-sense/src/ (NOT katgpt-core/src/sense — that dir does not exist; only sense_threat.rs lives in katgpt-core)
Game-theory in latent space — vector-op functors, coherence-driven re-estimation, zone-gated activation. Maps "stage"/"application"/"bypass"/"collapse" papers
riir-ai/crates/riir-engine/src/hla/
kernel.rs, forward.rs, types.rs — Higher-order Linear Attention (Transformer attention-layer replacement). Paper: Zhang et al. 2026. NOT the per-NPC belief (different layer, different repo).
Maps "attention layer"/"linear attention"/"recurrent state" papers to Transformer-scale ops
Generalized Stokes' Theorem substrate — d∘d=0 by construction. Maps "divergence"/"boundary flux"/"line integral"/"curl"/"Hodge"/"Fokker-Planck"/"mass conservation"/"manifold geometry"/"exterior calculus"/"Stokes". Curse-of-dim caveat: boundary-vs-volume wins only for d≤3 (maps, belief regions, KG embeddings) — NOT high-dim shards.
riir-infer/src/ (the riir-infer-core root) + crates/riir-infer-{gpu,laya}/src/
quant (quant/ EXL3-trellis lane, deltanet/ GDN lane, Bonsai2 rotation), loaders (gguf_loader.rs, safetensors_loader.rs, tokenizer.rs), transformer/ + simd/ forwards, the laya encoder lane (flat Vec<f32> ops over CPU + Metal MSL + CUDA kernels)
The substrate half of every serving claim — maps "quantization / weight format / kernel / encoder port / loader" papers; the G5-parity discipline (bit-identical port vs frozen captures, p-drift ≤ 1e-3) is the template for any forward port
Labeled decision serving with first-class abstention — maps "calibration / abstention / selective prediction / confidence / comparison arena" papers; published tables are CI-regenerated, never hand-typed, and G5 parity gates every laya number
A labeled (rule ↔ code-shape) binding corpus with measurable downstream accuracy — maps any "structured index / retrieval / compositionality / binding" paper; the withheld-pair OOD generalization axis is unmeasured almost everywhere
Adapter routing, KV compression, speculative decode = GOAT-tier framings. Latent-to-latent ops on belief/functor/shard/LatCal state = Super-GOAT-tier framings. Attempt Super-GOAT first. Defaulting to adapter routing when a latent reframe is stronger is the primary failure mode.
Consumer-surface miss (canonical failure #5 — arXiv:2608.29530 TPR, 2026-09-02): the paper's mechanism (role-filler binding, E = W(Σ fᵢ⊗rᵢ)+b) mapped onto the healer corpus (rule = role, code-shape = filler, oracle ground truth, measurable downstream top-1) MORE directly than onto any game surface — but the fusion hunt enumerated only game kernels: the factory table had no clippy rows, the MOAT table no clippy domain, and the repo-wide grep stopped at page 1 of matches, so the riir-rag/riir-refine wedge-consumption evidence was never read. The clippy fusion (riir-refine Issue 62: withheld-pair OOD bench + structured retrieval axis) surfaced only by user challenge — 5th rescue of the challenge-rescue class after NQF, LOPD, Le Critique, TTPO. Fixes applied: the priority ladder above, the clippy factory rows above, the step-4 consumer reframe, the MOAT clippy row, and the step-1 pagination rule.
Redirect to riir-train
Pre-check before ANY redirect: exhaust §3.5 modelless unblock paths first. A mechanism that looks training-only may be modelless-validable because its training-target MATH decomposes into shipped primitives (Flow Sampling is the canonical case: "trains a drift network via backprop" decomposed into dllm interpolant + Latent Field Steering reward-gradient + freeze/thaw replay buffer).
As of 2026-08-06: training is actively pursued, not lazily redirected. Applicable training papers (optimizer, loss, schedule, recipe, LoRA/OFT/IA3/QLoRA/DPO/GRPO/SFT/RL) get a Plan in riir-train/.plans/ with recipe + GPU-hours estimate + GOAT gate comparing trained-vs-modelless baseline. Only genuinely out-of-scope training (image-specific DiT we'll never train, medical imaging) gets the one-line redirect with explicit justification.
Runtime GRPO self-play stays in riir-ai (self-adaptive track — updates latent state, not weights). Model-based training (LoRA/SFT/GRPO on actual weights) lives in BOTH riir-train AND in-repo training pipelines (quest_grammar/grammar_training.rs in riir-ai, TernaryDraftModel.bits files in riir-ai/riir-refine). Quant-aware inference stays here; quant-aware training → riir-train. Never assume a repo is modelless-only — grep for existing training/model-based code first (pre-flight #5).
Workflow
Every path through this workflow terminates at §5 — the Claude verdict ping-pong — before anything commits. §1.5 / §1.55 / §1.6 grade the verdict; §5 is what makes it final.
0. Read & classify
Fetch via https://r.jina.ai/https://arxiv.org/pdf/{ID}. Ask: is the value in the training loop itself (optimizer/loss/schedule/RL), or in the math the training computes (closed-form drift, conditional score, Riemannian correction, steering formula)? Optimizer/loss/schedule → riir-train. Math → run §3.5 Path 0 first.
Label-anchoring hazard (canonical failure #2 — arXiv:2608.13335 NQF, 2026-08-18): the paper's own vocabulary ("Neural", "trained by gradient descent", "learning dynamics") is a routing hazard. Training-dynamics papers routinely carry closed-form modelless content — ignition-time formulas, closed-form trajectory families, initial-behavior limits, offline-computable spectra. NQF was routed wholesale to riir-train although its theorems carried a sigmoid-in-time adoption family, a t* = ln(1/ε)/ζ patience law, and a fresh-module leading-behavior predictor; surfaced only by user challenge (same rescue class as the LOPD canonical failure). Classification comes AFTER the §3.5 two-question decomposition + adversarial panel — never from the abstract's framing. The same hazard applies to NON-paper inputs (in-conversation proposals, model cards, issue texts): a 2026-10-07 encoder-lane proposal named "the Reflex decision family" as host and the note+POC were first routed to riir-reflex (public) although a trained-encoder serving lane is rethink's charter (private forever) — the family name was mistaken for the repo. The document's name for its host is routing INPUT, never the routing verdict; the defense is the family≠repo rule in the routing thumb + the MOAT rows + the §5 public/private tier check.
Test-time split hazard (canonical failure #4 — arXiv:2608.27448 TTPO, 2026-08-29): a "test-time training" paper carries TWO separable claims — the gradient objective (training track) AND the runtime decision layer the objective is built on (vote partition, asymmetric agree/penalize treatment, confidence scoring). TTPO's abstract framing routed the gradient objective to a riir-train plan filed as PRIMARY while the modelless decision layer — the half that fits the serving envelope — was downgraded to "recorded"; rescued by user challenge (4th rescue of this class after NQF, LOPD, Le Critique — user challenge is load-bearing, treat every challenge as a granularity audit). Fix: per-track verdicts (§1.5), discard-reason scrutiny (§3.5), serving-envelope fit (§3.5).
Three-track system — do NOT classify a repo as "modelless only" without checking pre-flight #5. The system runs three concurrent tracks: (a) modelless inference (primary hot path — no weights, BLAKE3-deterministic), (b) self-adaptive (runtime latent updates — self_evolve, EMA direction vectors, freeze/thaw cycles), (c) model-based (trained weights — ternary .bits files, LoRA/QLoRA adapters, Bonsai model comparison). A paper touching ANY track is actionable. The model-based track exists in repos OUTSIDE riir-train: quest_grammar/grammar_training.rs + quest_training.rs (LoRA rank=16-32, alpha=32-64, QLoRA 4-bit) in riir-ai; TernaryDraftModel with .bits weight files in riir-ai/riir-refine; self_evolve feature in riir-refine (shipped, G1-G7 ALL PASS). Before classifying a paper as "training-only, deferred" → grep for existing training code in the target repo (pre-flight #5).
Substrate ≠ value. Hardware/NMP/PIM/ASIC papers: the value is usually the technique (LUT INT→FP, shared ALU, sideband-tag) stripped of the hardware — grep simd_*, ternary, Plasma, Q4_K, LUT, from_bits for the software analog. Database/systems papers: riir-neuron-db IS a database (Pod + ShardIndex + Merkle + MAPE-K + Raven/δ-Mem + vibe KG) — the value is usually the access pattern. OS/kernel papers (io_uring, DPDK): substrate is Linux, value is usually the technique (lock-free queue, batching, zero-copy). Pure math/combinatorics: value may be a guaranteed-peak property on a per-entity scalar — see §1 step 4 game-context reframe.
Before PASS on hardware-vocab papers: (1) identify technique stripped of substrate, (2) grep for software-SIMD analog, (3) PASS only if both confirm no analog.
External code repo (not a paper)? Skip jina — §0.5 clones it into .raw/ for full-tree grep.
0.5 External-repo source access — clone into .raw/ (ephemeral, rm when done)
Papers distill fine over the wire (https://r.jina.ai/...). External CODE repos do not — distill verdicts, mining batches, and prior-art verification need full-tree grep, byte-accurate quotes, and multi-file reads that URL fetches cannot provide. When a research task needs an external repo's source:
bash
# Clone (shallow by default; full history ONLY if the task mines git log)
git clone --depth 1 https://github.com/<org>/<repo>.git .raw/<repo>
# Pin provenance BEFORE reading anything into a verdict:
git -C .raw/<repo> rev-parse HEAD # record this sha + the license in the note
Rules (all mandatory):
Location:.raw/ at the root of the repo whose .research//.plans/ will hold the output (e.g. katgpt-rs/.raw/ds4/). .raw is ephemeral scratch — gitignored in every workspace repo, never committed, never a sibling repo.
Pin before verdict: record @ <full-sha> + license (Apache-2.0/MIT/...) in the research note BEFORE deleting the clone — the sha is the provenance; the clone is throwaway. Every quote later cited must match the pinned sha (precedent: riir-refine Batches 96–98 distills cite antirez/ds4 @ 110afdd8… with quotes "re-verified at mine time").
Read-only. Do not modify files inside .raw. If a patch experiment is needed, copy the file out first; never let experiments dirty the clone mid-verdict.
No build-graph contamination: never add a Cargo/path dep into .raw (BOUNDARY.md + standalone-dep gates), and never include .raw in codebase/pre-flight grep scopes — it is NOT shipped substrate, so grepping it inflates hit counts and makes a mining candidate read as "already ships".
Cleanup is part of the task: when the verdict/note/plan is committed, rm -rf .raw/<repo> (or the whole .raw/). A finished research task with a live .raw/ entry is an unfinished task. If a session dies mid-task, treat any stale .raw/ entry as UNTRUSTED on resume — delete and re-clone at the pinned sha; never verdict against a clone whose commit you did not pin yourself.
Shallow unless history-mining:--depth 1 by default; full clone only when the task needs git log/diff archaeology (fix-pair mining, regression hunts) — and say so in the note. For very large repos where only a few files are read, --filter=blob:none fetches blobs on demand.
1. Distill fundamentally — fuse, don't direct-map
Find the transferable primitive (the geometric/spectral/information-theoretic insight that works without the paper's training setup). Then look for fusion opportunities: cross-pollinate with existing notes/plans/shipped primitives to synthesize something novel.
Fusion protocol:
Grep ALL workspace repos, BOTH layers (notes AND code), in parallel via subagents — the product set of 7 first, then the wider set (derive the set, never hand-type; see §Repos — 15 enumerated today). Do NOT stop after the first repo or layer, do NOT wait for user prompts to grep the next repo. Grep results are PAGINATED — a full first page is NOT a finished sweep: run offset until exhausted, or scope the grep per-repo so every repo's hits are visible. Closest cousin is frequently in the OTHER repo (cross-repo fusion) or in CODE not notes (mechanisms ship without research notes — evolve_belief is a per-NPC recurrent belief-state kernel with no .research/ framing). Grep:
Every .research/ + .plans/ folder in the workspace (derive both — typed counts of either rot within weeks; riir-train included: applicable training papers get Plans per §3.5 Path 0.5)
riir-ai/.docs/ (the moat/selling-point book — grep alongside .research/ so you don't claim novelty over a pillar that ships)
All shipped src//crates/ trees (notes describe intent; code describes what shipped)
The twelve Super-GOAT factory modules above
Vocabulary translation BEFORE grepping. Papers and codebase use different words. List the paper's 3–5 key terms; for EACH, brainstorm ≥2 codebase equivalents ("if we shipped this, what would we call it?"). Grep BOTH sets. read_filevocab.md (sibling file) for 8 standing vocabulary tables + worked examples. Coefficient-shape mechanisms (adaptive blend / mixture coefficient / interpolation ratio / ensemble weight): shipped blends hide as hand-tuned constants (W_EVO = 0.6, alpha: f32, trust: f32), never under the word "blend" — use vocab.md §6's grep set and run it across ALL repos, never scoped to the expected home (canonical failure #3).
Latent-space reframing BEFORE verdict. Re-cast the core mechanism as a latent-to-latent op on the codebase's kernels: (a) per-NPC belief, (b) latent_functor/ ops, (c) cgsp_runtime/ curiosity, (d) LatCal fixed-point commitment, (e) NeuronShard style_weights/dendritic/MerkleFrozenEnvelope/Raven/AnyRAG, (f) DEC Stokes operators (d, δ, hodge_decompose, DecFlowField). If you only reach adapter routing/KV/spec-decode framing → likely GOAT, missed the Super-GOAT angle.
Game-context AND consumer-context reframing BEFORE verdict (especially when step 3 returns no hits). Game question: how does this mechanism manifest as a per-NPC behavior signal / crowd pattern / selling point in MMORPG context? A guarantee → what per-entity scalar does it bound, can that drive behavior (salience cadence, curiosity trigger, consolidation window)? A combinatorial structure → does it appear in NPC routines, market cycles, quest scheduling? A number-theoretic property → fairness/diversity/coverage on a game signal? A geometric property → NPC phase scheduling, spatial spread, fog-of-war coverage? Consumer question (priority #2, same evidence bar): does the mechanism manifest on a healer surface — (a) the retrieval index / corpus structure, (b) fix-trajectory memory / selection, (c) the score-bench / eval axis, (d) drafter-pruner selection, (e) rustc_errors playbooks? One sentence each; a bare "no" without reading the surface is the canonical-failure-#5 shape. If either reframe is missing, you are NOT ready to PASS.
Zero hits ≠ novelty. Most likely you're using the wrong vocab — try a third semantic angle (grep output behavior "swap when X" instead of mechanism name "tightness monitor") before claiming "no prior art".
List the 2–3 closest cousins across all repos. Ask: what novel combination of paper × A × B produces a capability none has alone? Write that into the note's §Distillation as a Fusion subsection even if unplanned.
Verdict by tiers (§1.5). Create research .md at the right repo. Naming: {NNN}_{Short_Title}.md (next free number, monotonic, never reused). Format: read_filetemplates.md (sibling) — canonical example: katgpt-rs/.research/238_LoRA_Muon_Spectral_Low_Rank_Manifold.md.
1.5. Novelty gate — is this Super-GOAT?
Score all four:
No prior art? Grep notes + plans + shipped code across all repos, using BOTH paper vocab AND codebase-vocab alternatives. Read every grep hit's TL;DR before claiming novelty — a filename match is a lead, not confirmation. Symmetric rule for the opposite direction: a cousin NAME match is a lead toward coverage, never coverage itself — "already ships" claims require the §3.6 signal-diff check. Mechanisms ship under non-obvious names (DiPOD's "interleave self-distillation when ELBO drifts" ships as "coherence-driven re-estimation scheduler when coherence < tau_reest"). When the candidate selling point touches per-NPC + memory + personality + swap, the riir-ai/.research/ corpus is saturated — grep it broadly and READ every hit before claiming novelty.
New behavior class? Not better numbers — a capability no incumbent has. Requires step 4 game-context reframe to detect. If you reached Q2 without step 4, go back.
Product selling point? Can you finish: "Our NPCs/systems do X that no competitor can"? If not → Gain.
Force multiplier? Connects to ≥2 existing pillars/systems. Solo novelty = GOAT, not Super-GOAT.
All 4 YES → Super-GOAT. Mandatory outputs (in this same session):
Open primitive → katgpt-rs (generic math, no game semantics).
Architectural GUIDE → private selling-point doc. Repo by selling-point domain: riir-ai/.research/ for game-runtime/belief/functor/self-learn; riir-chain/.research/ for chain/LatCal/commitment/sync-bridge; riir-neuron-db/.research/ for shard/freeze/consolidation/AnyRAG/vibe/Merkle. Cross-domain selling points → primary guide in the repo owning the boundary being crossed. Guide includes: TL;DR with commercial value, distilled modelless primitive, connection map, latent↔raw boundary, what's private vs open, validation protocol, P0–P3 priority.
Plan(s) → appropriate .plans/ folders.
No "candidate" escape hatch. Writing "all 4 YES" / "passes the gate" / "Super-GOAT candidate" anywhere in a note triggers the mandatory outputs in THIS session. If not confident on all 4 → write "fusion idea, novelty TBD" + create .issues/ entry. "Candidate" is not a deferred-commitment escape — it either triggers the guide now, or downgrades to an issue.
One verdict per track, not one per paper (TTPO lesson, 2026-08-29). When a paper yields BOTH a training recipe AND modelless extractions (the Path 0 inventory rows), score this gate SEPARATELY per track. A training-track plan and a modelless-track plan are two claims with independent tiers — a Q1 kill on one does NOT cascade to the other. "Recorded in the note" is NOT a tier (same violation class as the banned "candidate" escape hatch): every Path-0 inventory row ends in exactly one of (a) filed plan (GOAT+), (b) .issues/ entry (fusion idea / novelty TBD), or (c) an audited discard whose reason cites the mechanism-level §3.6 diff — see the discard-scrutiny rule in §3.5.
Pin the claim BEFORE searching it (§4 precondition — TTPO lesson, 2026-08-29). Prior-art search 3 below is only meaningful against a pinned claim. Before running ANY §4 search, write the claim in one sentence: "<mechanism> for <surface/consumer>, consuming <signal>, distinguished from <closest cousin> by <delta>." A novelty check against an unpinned claim is void — class-level prior art (JitRL/OptPO killed "training-free test-time policy optimization" the CLASS, not the asymmetric-partition + calibration-law MECHANISM) will silently kill anything, and the verdict regresses to vibes. If you cannot write the sentence, you have not distilled enough to claim novelty either way.
1.55. PASS vs Gain — no middle tier
PASS = no new research/plan files. Gain = files. Scan the paper for actionable improvements before verdict.
Ships + actionable improvements → Gain (create .issues/ per AGENTS.md "Create issue at .issues for poc, proof, optimization or refactor task, do not create plan", and/or plan behind feature flag).
Ships + no actionable → Pass (verdict + one-line reason + closest shipped cousin, in conversation only; no files).
Doesn't ship + modelless → Gain or higher.
Doesn't ship + training-only → run §3.5 Path 0.5. Applicable → Plan in riir-train/.plans/. Out of scope → one-line redirect with explicit justification.
PASS-Redirects line (mandatory): PASS verdicts still update the 1–3 closest shipped cousin .research/ notes with a one-line reference. Format:
Must include arxiv ID AND full title (so grep arxiv:ID AND grep "Title" hit). Prevents paper-number invisibility — without this, a future session greps the arxiv ID, finds nothing, re-distills from scratch.
Actionable = Gain. Actionable: paper data contradicts a current config default; exposes a failure mode with no existing mitigation; unblocks a deferred task. NOT actionable: "validates our design" / "theoretical lens" / "could inform a future config". Unsure → not actionable → Pass.
Reverse-grep before PASS (mandatory). Before any PASS, grep the codebase for documented gaps the paper could fill: .docs/ for Limitation|deferred|TODO|FIXME|gap|pending; .benchmarks/ for Caveat|deferred|artifact; .rs comments for TODO|FIXME|deferred|limitation near the paper's vocab. If ANY hit maps → Gain, not Pass. Compact heuristic: "Is there any documented limitation this paper could fix?" — if you can't answer "no, I checked" with evidence, don't PASS.
Training papers third defense: both defenses above can return clean and still produce a false PASS if "model-based track" is narrowed to one pipeline or one repo. Before PASS-ing any training paper: (1) read_file riir-train/.docs/02_pipelines/training_data_pipeline.md and ask: does a training pipeline for this domain already exist? (RL, distillation, linear attention, trajectory collection). (2) grep ALL repos for existing training/model-based code (pre-flight #5) — training pipelines ship OUTSIDE riir-train: quest_grammar/grammar_training.rs + quest_training.rs (LoRA training in riir-ai), TernaryDraftModel (ternary weights in riir-ai/riir-refine), self_evolve (self-adaptive in riir-refine). If ANY existing training/model-based code maps to the paper's domain → Gain, not Pass. Canonical failure (arXiv:2511.18538): Code LLM survey falsely PASSed because agent checked only riir-train (game AI pipeline), missing quest_grammar's code-domain LoRA training + riir-refine's ternary drafter + Issue 010 (generalization gap that RLVR would fix).
1.6. MOAT gate per domain + promote/demote
Tier verdict measures how strong. MOAT gate measures whether it strengthens THIS repo's moat. Mismatch → reroute.
Domain
MOAT bar
In scope
Out of scope
katgpt-rs
Fundamental/principle/base primitive via fusion; promote/demote tracked per stack
Private-forever moat. The trained-encoder decision product — nothing here may land in a public repo
The trained-encoder serving rung, its weight artifacts and serving posture, escalation-lane enforcement + rate guard, its economy planes
The escalation GRAMMAR → instinct (public); the encoder forward SUBSTRATE → riir-infer (public); re-fit recipes → riir-train; ANY public-repo file — this tier's product direction never lands in katgpt-rs/riir-infer/riir-reflex/riir-instinct/katgpt-web
Ledger bound-enforcement stays riir-dapps (the agent only proposes); generic rating/Beta math stays katgpt-core
9 sloppy-test pillars live in riir-ai/.docs/03_pillars/README.md — read_file03_pillars/README.md + 04_supergoat_candidates/README.md before any "does this become a pillar?" verdict. Sloppy test: if it doesn't exist, the system goes structurally sloppy — not slower, broken.
katgpt-rs per-stack ledger: every primitive gets a feature flag + benchmark + GOAT gate; the verdict note records which stack slot (attention/KV/sampling/speculative/pruning) + promote-to-default or stay-opt-in. Re-gate on feature touch. Demote the loser when a newer primitive wins the same slot.
Show full SKILL.md (3,594 more words)Show less
1.7. Pre-plan cherry-pick audit
If your plan will consume/wire/fuse with a katgpt-rs primitive into riir-* → run the goat-audit skill before opening the plan. Catches (a) stalls (default-on in katgpt-rs ≥7 days with zero riir-* consumer), (b) DRY violations (riir-* shipping a local duplicate of substrate — e.g. defining its own KVCache instead of consuming katgpt-transformer). When NOT to run: purely katgpt-rs-internal; bug fix no cross-repo angle; training-only genuinely out of scope.
2. If gain/GOAT, plan it
Plan .md in katgpt-rs/.plans/ (modelless), riir-ai/.plans/ (runtime/game), riir-chain/.plans/ (chain/LatCal), riir-neuron-db/.plans/ (shards), riir-train/.plans/ (training-efficiency per §3.5 Path 0.5). ## Phase N sections, - [ ] per task (- [x] done, - [-] deferred). Planning into riir-train IS allowed + actively encouraged for applicable training-efficiency papers. Super-GOAT plans come AFTER the guide — guide is strategy, plan is execution.
Plan format + GOAT gate rule + UQ "Report the Floor" extension:read_filetemplates.md. Compact: every new technique needs feature flag + benchmark proving gain before promoting to default; demote loser if newer wins same slot. UQ-bearing primitives (distributions/intervals/coverage) MUST benchmark against conformal-naive floor (ConformalIntervalCalibrator<SeasonalNaiveForecaster>) on CRPS/coverage/Winkler — if they can't beat it, the gate FAILS.
3. Implement to unblock
If a plan is blocked by a missing primitive, implement the minimal version. After GOAT check + proof of gain: promote to default if it wins, demote the loser.
3.5. Modelless unblock — MANDATORY before any riir-train deferral
Before deferring ANY gate/task/mechanism to riir-train, exhaust modelless paths first.
Path 0 — training-target decomposition. Is the value the training-loop innovation (new optimizer/loss/curriculum/RL) or the math it computes (closed-form drift, conditional score, Riemannian correction, steering formula, regression target)? If the math → decompose into components, grep for modelless analog of EACH. ALL have analogs → MODELLESS-VALIDABLE (no deferral). SOME missing → check paths 1–3 for those.
Path 0 asks TWO questions per component, not one. (1) Coverage — does an analog already ship? (feeds the §1.5 Q1 novelty gate). (2) Extraction — can the component be computed WITHOUT gradient descent: closed-form solution, derived schedule, predictor, bound, ordering law, limit behavior (zero-init / mean-field / leading-order), offline-computable spectrum? A row with no coverage but yes-extraction is the STRONGEST finding class (open-primitive candidate → katgpt-rs) — marking it "no analog" and moving on is the canonical failure #2 miss (the NQF Path-0 table said "no closed-form plateau predictor" and stopped, instead of noticing the paper HAD one to extract).
Path 0 outputs an INVENTORY, not just a verdict. Complete the component table even when the funnel lands on Path 0.5 — every row marked "analog exists" or "partial" is a candidate open primitive. Before dismissing any such row as "covered", run the signal-diff check (§3.6): read the shipped analog's core formula and state what signal it consumes vs the paper component's. Name-level cousin match ≠ coverage — the diff that surfaces a real gap sounds like "relevance vs utility" / "history vs current-query" / "aggregate vs counterfactual", and it is answerable with ONE read of the cousin's code. Canonical failure (2026-08-16, arXiv:2608.13040 LOPD): the modelless δ (counterfactual with-vs-without utility gate) was dismissed as "covered by EvidenceTier / engram gate" — but engram's kernel gate is σ(dot(q,k)/τ) (relevance-only; the shipped zero-query test pins gate=0.5 fusion of useless slots as correct behavior) and EvidenceTier is history-only 3-tier, while δ is query-conditional utility at use time. Reverse-grep could NOT catch this class — the gap was pinned as a feature, not a TODO (→ katgpt-rs Issue 656; surfaced only by user challenge, one grep into kernel.rs had sufficed).
Granularity rule (canonical failure #3 — arXiv:2608.16739 Le Critique, 2026-08-20): diff the CONTAINING expression, not only the named cousin. The BetaPosterior↔TETHER signal-diff was run and correctly returned “different shapes” (per-candidate posterior vs global mixture weight) — but the formula that CONSUMES BetaPosterior, W_EVO·evolution + W_RATE·reliability in riir-refine select_best_candidate, is exactly TETHER's b(ρ) = (1−ρ)p1 + ρp2 with ρ hand-pinned at 0.4 and never swept, while the realized-outcome feed (EvolveRecorder::record_outcome) already fires on every applied fix (→ riir-refine Issue 033; found by user challenge — third rescue of this class after LOPD and NQF; user challenge is a load-bearing defense layer, treat every challenge as a signal-diff granularity audit). Rule: after diffing cousin C, read ONE level up — the expression combining C with other signals. Hand-tuned constants composing C (const W_X: f32 = 0.6) are the paper's shape wearing a different name; the grep is W_[A-Z]|const [A-Z_]+: f32 on C's caller (one cross-repo grep had sufficed).
Three-track adversarial panel — MANDATORY whenever the first-pass classification touches training (routes to riir-train, PASSes a training-adjacent paper, or the abstract carries optimizer/loss/RL/backprop framing). Spawn 2 spawn_agents IN THE SAME parallel batch as the §4 prior-art searches (one spawn round total). Advocacy briefs — neutral merging is the coordinator's job; each agent argues FOR its track and does NOT inherit your classification:
No-GD advocate (constraint #1 tracks a+b: modelless inference + self-adaptive runtime): "Extract every closed-form, derived quantity, predictor, schedule, bound, ordering law, limit-behavior, or offline-computable spectrum in this paper. For each: can it ship without training? Which repo/module? What GOAT gate? Argue FOR feasibility. Do NOT judge whether it already ships — another pass owns coverage."
Model-based advocate (constraint #1 track c: trained weights): "Extract every actionable training-recipe item (optimizer/loss/schedule/init/rank/width/capacity/data). Which existing pipeline does it improve (riir-train Plans, quest_grammar, TernaryDraftModel, edge_lora, riir-infer's EXL3/GDN quant lanes)? Recipe + GPU-hours + GOAT gate vs modelless baseline. Argue FOR applicability."
Coordinator merges both into the Path 0 table. Discarding an advocate finding requires a one-line auditable reason in the note (extends the §3.6 signal-diff discipline to track routing — the coordinator can still be anchored; this is the backstop). A discard reason must survive mechanism-level scrutiny (TTPO lesson, 2026-08-29): class-level prior art ("JitRL ships training-free TPO"), pure cost notes ("2× decode"), and action-level coverage ("the prune action ships in Plan 133") are NOT sufficient kills when the finding includes a calibration law, a product-form signal, or a wiring fusion those references do not carry. Cite the exact uncovered delta, or do not discard. Brief hygiene — do not leak classifications through the stack description. Describe repos by WHAT SHIPS (file + type names: “riir-refine self_evolve.rs ships select_best_candidate + EvolveRecorder outcome recording”), never by your shape conclusion (“BetaPosterior candidate selection” pre-answers the examination the advocate should perform). The Le Critique miss propagated exactly this way: the No-GD advocate inherited the framing and never read the selection formula — its only riir-refine blend proposal was the retrieval-layer one the framing made salient. Clearly inference-side papers (KV cache, sampling, spec-decode) skip the model-based advocate; clearly pure-hardware papers skip the panel only because the §0 technique-stripping pass already ran.
Path 0.5 — training-cost-weighted re-evaluation (DEFAULT for training-efficiency papers as of 2026-08-06). Training efficiency is actively pursued. Applicable training paper → Plan in riir-train/.plans/ with recipe + GPU-hours estimate + GOAT gate comparing trained-vs-modelless baseline. Only genuinely out-of-scope → one-line redirect with justification.
Track priority — serving-envelope fit (TTPO lesson, 2026-08-29: "test-time gain is modelless shallow reasoning"). When BOTH tracks produce filed plans, rank them by whether the mechanism's decision layer runs INSIDE the stack's hot path: modelless shallow selection (branch rollouts, EqR/pruner/selector choices, healer spans, 20 Hz game ticks) vs outside it (gradient TTT / per-item training — structurally unaffordable where base competence ≈ 0 or per-fix budget ≫ serving budget; canonical datum: the L4 fixer's ~191 s/fix + 0/60 EM). The envelope-fit track is the PRIMARY plan; the other is SECONDARY and says so in its Status line. GPU-hours affordability (Path 0.5) is necessary but not sufficient — a cheap training run that never reaches a production consumer loses to a modelless gate on the live serving path.
The model-based track = ALL training pipelines across ALL repos, not just riir-train (GDN-blog lesson + arXiv:2511.18538 lesson). read_file riir-train/.docs/02_pipelines/training_data_pipeline.md before PASS-ing any training paper — RL/distillation/linear-attention/trajectory-collection pipelines all count. BUT ALSO grep non-riir-train repos for training code (pre-flight #5): quest_grammar/grammar_training.rs + quest_training.rs in riir-ai (LoRA training for quest grammar + quest corpus), TernaryDraftModel in riir-ai/riir-refine (ternary-weighted drafter loaded from .bits files), self_evolve in riir-refine (self-adaptive latent-first loop). Training is NEVER "deferred indefinitely" — it is actively pursued across the whole stack.
Systematic backstop: when ≥3 training-recipe gaps accumulate from PASS-Redirects across repos, batch them into a single riir-train/.plans/NNN_training_recipe_gap_backlog.md plan.
Three modelless unblock paths (check ALL before deferring, AFTER Path 0 fails; run Path 0.5 AFTER all three fail):
Freeze/thaw snapshot correction (MerkleFrozenEnvelope) — can a corrected snapshot, thawed at inference, fix a systematic bias (doubled signal, position offset, attention asymmetry)?
Raw/lora reader-writer hot-swap (LoraPair { reader, writer }; LoRAHotSwap, dispatch_lora_merge in riir-ai) — can a deterministically constructed (not trained) adapter fix it? Closed-form (scale-by-0.5, zero-out-positions, identity-minus-projection) rather than gradient descent?
Latent-space correction (dot-product + sigmoid gate, per constraint #2) — project latent onto correction direction, gate the output. Modelless analog of trained adapter.
Decision protocol:
Paper appears to need training
→ Path 0: value = MATH not training loop? YES → decompose + grep modelless analog of each
ALL have analogs + a use case → MODELLESS-VALIDABLE. No deferral.
SOME missing → paths 1–3 for those. ALL fail → Path 0.5.
(either way: rows marked exists/partial → signal-diff EACH before "covered" — §3.6)
→ Systematic characterizable cause? YES → paths 1→2→3. ALL fail → Path 0.5.
→ Path 0.5: applicable → Plan in riir-train. Out of scope → redirect with WHY.
Documentation requirement for every riir-train Plan: Path 0 decomposition (which components had analogs, which didn't) · paths 1–3 checked + why each failed · what specifically requires GD no deterministic construction can provide · Path 0.5 affordability (recipe + GPU-hours + GOAT gate) · dual-track contribution (modelless inference primitive + trained weight artifact).
3.6. Defend-wrong PoC — MANDATORY before any "already ships" / "parity" verdict
Before claiming a mechanism "already ships", achieves "parity", or "covers" the paper's loop, distinguish three claim types:
Claim
Proof
Architectural ("runtime analog exists")
grep + read code (sufficient)
Latency/resource ("modelless, sub-µs, no GD")
criterion bench
Quality ("matches/beats paper's numbers")
head-to-head PoC on controlled toy — architectural reasoning NOT sufficient
Failure mode: claiming all three with only architectural evidence. Grep proves mechanism exists; it does NOT prove it performs as well as the paper's version.
PoC mandatory when: quality-parity verdict ("matches", "competitive with", "covers at parity"); qualitative Super-GOAT/GOAT claim; any PASS that downgrades on grounds "the runtime analog already ships" (architectural-only PASS is the #1 false-PASS mode).
Scope extension (2026-08-16, LOPD lesson): coverage-dismissals inside Gain verdicts need the same defense. Claiming a paper's secondary mechanism is "covered by cousin X" — while filing the primary mechanism as a training plan — is a verdict, and it must rest on mechanism-level evidence, not a name match. The proportionate defense is the signal-diff check: read the cousin's core formula/gate and state what signal it consumes (relevance? history? aggregate?) vs the paper component's (utility? current-query? counterfactual?). Full PoC remains mandatory only for quality-parity claims; a dismissal needs one read of the cousin's formula. Two blind spots this closes: (a) every other defense (§3.6 triggers, reverse-grep) fires only on PASS — a Gain verdict's secondary dismissal had NO guard; (b) reverse-grep finds documented gaps (TODO|FIXME|limitation) — this miss class ships undocumented, with the blind spot pinned as correct behavior in the cousin's own test (engram zero-query gate=0.5 fusion; → katgpt-rs Issue 656).
PoC NOT required when: pure architectural redirects (no quality claim); training-only genuinely out of scope; latency-only claims (single bench suffices); low-confidence verdicts that explicitly mark quality claim unproven + create .issues/ follow-up.
Where the PoC lives:riir-ai/crates/riir-poc/ (defend-wrong R&D crate). Three competitors minimum: paper's mechanism (or distilled modelless analog), frozen/no-adaptation baseline, shipped runtime analog. Head-to-head on controlled toy domain, no training, print verdict table. CARGO_TARGET_DIR=/tmp/... + clean up.
PoC defends OR refutes. If it refutes quality: do NOT silently revise. Record raw numbers in research note §"PoC Addendum". State which axes were confirmed (architectural, latency) vs refuted (quality). Verdict stands on confirmed axes; refuted axis becomes tracked follow-up. PoC stays as permanent regression check.
4. Published prior-art search — MANDATORY before any novelty verdict
Hard gate. Not optional. Before claiming ANY novelty (Q1), web-search for published prior art on the headline technique. The canonical failure mode: claiming "ternary MoE is novel" when one search for "mixture of ternary experts" would have found published prior art. The miss happens because the agent only read the paper's own references + grepped the codebase, never searching broader literature.
Mandatory searches (run ALL before any novelty claim):
Headline technique verbatim ("mixture of ternary experts", "BitNet distillation recipe").
2–3 component techniques.
The selling-point framing you're about to claim.
Recent (2-year) surveys on the topic — they name the competitive landscape.
If any published paper does what you claim is novel → downgrade Q1 BEFORE writing the verdict. Cite explicitly. Do NOT write the verdict then discover prior art later.
Parallelize via subagents: 2–3 spawn_agent in parallel (headline search, component search, codebase grep). Web catches published prior art; codebase grep catches shipped prior art. Both mandatory; neither substitutes.
Re-run after corrections: discovering prior art in a later pass is NOT complete until you've (a) updated novelty verdict, (b) re-checked surviving novelty claim, (c) committed the correction. Don't leave notes in overclaimed state.
4.5. Optional deeper search
If §4 surfaces rich landscape, use web search for deeper exploration of specific papers/authors/follow-ups. Not mandatory, valuable when prior-art landscape is dense.
5. Final verdict — Claude ping-pong (MANDATORY before any commit)
A verdict is not DONE until the Claude reviewer has seen it. The gates above (§1.5 / §1.55 / §1.6) are your self-grade. Before anything commits, negotiate the verdict against the Claude reviewer sub-agent. This is the research-side merge gate and it operationalizes the standing owner rule — ask Claude for verdict and make decision for any owner gated. A self-graded verdict that never met the reviewer is the research analogue of an unreviewed merge.
Availability premise (measured, not assumed): the reviewer rides the harness tool request_verdict (the claude-sub-agent-verdict reviewer; round cap agent.verdict_max_rounds, default 3) — a harness capability rather than repo substrate, so it is cited to a WORKED RUN, not a spec document: the executed precedent is mmorpg-editor Proposal 007 (on disk seal-game-editor/ — the contract-name rule), which ran this exact gate through two #Verdict: rounds to the cap-3 contract on the claude backend, recorded in its Status + the plan/issue trail. A harness without the tool does NOT skip the gate — it takes the sub-agent path below, a first-class route to the same binding verdict, never a degraded one. The note records which path ran.
Protocol (the tool enforces the mechanics; the Summary is YOUR job):
Compose the ## Summary — complete, because the reviewer CANNOT see your context. It carries: paper ID + full title · the pinned one-sentence novelty claim (§1.5's precondition form) · per-track verdicts with tiers (never pooled) · the Path 0 inventory outcome (analog / extracted / deferred per component) · routing (repo + every file created or edited, exact {NNN}_{Short_Title}.md names) · the 2–3 closest cousins considered and why they don't kill the claim · and the weakest point you know of, named by you — the reviewer finds it anyway, and naming it first costs one round instead of two.
Round 1 — call request_verdict with the Summary + the note's file paths (so the reviewer can read the note itself) + the literal instruction: "Reply with a verdict that MUST start with #Verdict: AGREE or #Verdict: REVISE followed by bullet-point reasons."
REVISE → address EVERY bullet with EVIDENCE — a grep result, a verbatim quote at the pinned sha, a bench number, a file path — then call again with the SAME session_id (the negotiation stays in one thread). Rephrasing your position without new evidence does not address a reason.
AGREE (and you agree) → state your own agreement, restate the final agreed ## Summary, pass final_round: true on that closing call, and only then stage the NAMED files + commit + push.
Round cap (default 3) → the tool refuses further calls. STOP. Present the remaining disagreement to the USER with both positions — the user is the owner gate, never a tiebreak you award yourself. A capped disagreement escalates; it does not self-resolve by picking your side.
Scope — mandatory vs optional:
Mandatory: Super-GOAT and GOAT verdicts · every riir-train Plan filing (Path 0.5) · every verdict that creates or rewrites a .research/ note, .plans/ file, or architectural guide · every advocate-finding discard (§3.5) · every "already ships"/parity claim granted without the PoC (§3.6 exemptions) · every owner-gated call (feature promotion/demotion, default-on flips, §1.6 tier re-routing).
Optional (encouraged): a plain PASS whose §4 sweep found zero redirect-worthy cousins — a checkable claim, since the searches are in the Summary and auditable next round; a routing call the owner already made in-conversation this session.
What the reviewer checks — hand it the hooks: novelty claim vs the §4 searches actually cited (never asserted) · §3.6 signal-diffs on every coverage dismissal · routing vs the MOAT table + fusion priority ladder · public/private tier check: every file named for creation is classified public-safe vs private-tier product direction — a public-repo destination (katgpt-rs, riir-infer, riir-reflex, riir-instinct, katgpt-web) carrying rethink-class content (serving arms, vessel/hosted-only posture, ESC enforcement, economy planes) is a REVISE; default private on any doubt · per-track separation (no cascade from a training-track kill) · discard reasons surviving mechanism-level scrutiny (§3.5) · file hygiene (numbers from .highwater, PASS-Redirects lines present, no "candidate" escape-hatch wording).
Standing failure modes (how this gate dies quietly — pre-registered, not yet measured):
Self-AGREE drift — summarizing your way to AGREE by omitting the weak axis. The Summary must name the weakest point explicitly; an omission voids the verdict, and a re-opened verdict after files shipped costs a renumber + citation rewrite (the exact cost the numbering rules exist to price).
Cap-racing — burning rounds re-arguing instead of producing evidence. Each REVISE round must ADD something checkable; a round that adds nothing is a round the user now has to arbitrate.
Post-AGREE drift — editing a verdict-bearing section after AGREE without re-running the gate. A material post-AGREE edit re-opens the negotiation on the SAME session_id before the next commit; the commit after a final_round close is byte-frozen to what was agreed.
Reviewer-unavailable — no request_verdict in the toolset? Run the SAME negotiation through a generic sub-agent reviewer (spawn_agent in this harness; Agent/SendMessage in Claude Code — the capability is a fresh reviewer that cannot see your context, whatever the local spelling) with the same Summary + the same #Verdict: reply contract. It is an equal verdict, not a weaker one — record which path ran, and never label a fallback verdict provisional (a permanently-provisional gate is an unexecutable mandatory, this file's own most-repeated shape).
Constraints (non-negotiable)
Three-track system (not modelless-only) — the stack runs three concurrent tracks: (a) modelless inference (primary hot path — no weights, BLAKE3-deterministic, zero-alloc), (b) self-adaptive (runtime latent updates — self_evolve, EMA direction vectors, freeze/thaw cycles, CommittedFieldBlend — NO base weight mutation), (c) model-based (trained weights — ternary .bits files via TernaryDraftModel, LoRA/QLoRA adapters, Bonsai model comparison, quest_grammar training pipelines; the quant-format/kernel substrate for trained weights ships public in riir-infer). Do NOT characterize any repo as "modelless only" without checking pre-flight #5. Closest to "training" for the modelless track: freeze/thaw cycles, raw/lora hot-swap with deterministically constructed adapters (not trained), latent direction-vector updates at runtime. Before any riir-train deferral, exhaust §3.5. AND check for existing model-based code in the target repo.
Latent-to-latent preferred — operate in latent space as long as possible. Decode/project only at boundary. Sigmoid, never softmax, for projections onto learned directions. Semantic (emotion/mood/curiosity/style) → latent. Physical (position/HP/wallet) → raw, deterministic, synced.
Freeze/thaw over fine-tuning — only runtime weight mutation is swapping a frozen snapshot (atomic, versioned, BLAKE3-checked) or applying a deterministically-constructed LoRA overlay (raw/lora hot-swap, no GD). Never mutate weights in-place during inference. Gradient updates (after §3.5) → riir-train.
Self-learn / adaptive CoT welcome — runtime curiosity, latent prediction, trajectory folding, collapse detection. Update latent state / direction vectors / routing tables, NOT base weights.
7-repo discipline (the product/distillation set; canonical list in katgpt-rs/AGENTS.md §"Repo count") — katgpt-rs (public) → riir-ai → riir-chain → riir-neuron-db → riir-train (all private) + riir-game-sdk (facade) + riir-dapps (dApp layer: game outcome → generic chain settlement, added 2026-08-20). It read "8-repo" and included riir-armageddon until 2026-09-03; that repo was retired 2026-09-02 (owner act, directory gone). The WORKSPACE around that set is 27 contract repos at the 2026-10-07 read (canonical: katgpt-rs/AGENTS.md §"Repo count"; the Instinct/Rethink carve added two 2026-10-03) — riir-refine, riir-auth, riir-dao, riir-viewbridge, riir-rethink carry .research/ too (healer-domain notes file in riir-refine per the MOAT table; rethink's carries the ESC-depth line; instinct's is created on first use); riir-infer, riir-reflex AND riir-instinct are PUBLIC first-class routing targets (MOAT rows in §1.6 — inference-substrate papers file in riir-infer, decision-serving papers in riir-reflex, open-stack/grammar papers in riir-instinct); the rest (riir-kat, riir-deployer, riir-esp32, katgpt-web, riir-mmorpg-examples, riir-shader, riir-llm, riir-reflexer) are consumers/protocol/leaf/POC/presentation surfaces — routing targets of last resort. Read a count in prose as a claim, not a fact — derive the set (§Repos). Training how never leaks to katgpt-rs; chain IP in riir-chain; shard IP in riir-neuron-db; SDK stays facade over riir-games-shared; riir-infer + riir-reflex are the sanctioned public substrate/serving tiers, not runtime/chain/shard/train IP.
SOLID, DRY — per katgpt-rs/.contexts/optimization.md. Zero-alloc hot paths. Pre-computed lookup tables. Fixed-size arrays for bounded domains.
Tests/examples — before/after showing the gain. Latent ops: projection preserves ranking. Freeze/thaw: readers never see torn snapshots.
Plasma → Hot → Warm → Cold → Freeze tiering — perf on game side (plasma/hot budget), security on chain side (cold/freeze commitment, BLAKE3-hashed, tamper-evident). Latent state crossing sync boundary MUST be raw scalars (valence/arousal/desperation/calm/fear), never the full embedding.
Latent vs raw space rules (critical for game AI)
Physical (position, velocity, HP, wallet): MUST be raw exact. Deterministic replay, quorum sync, anti-cheat require bit-identical reconstruction.
Semantic (emotion, mood, curiosity, style, habit): SHOULD be latent via dot-product + sigmoid onto learned direction vectors.
Social (encounters, relationships, factions): SHOULD produce KG triples from latent/embedding proximity, not raw coordinate distance.
Sync boundary: if data flows through SyncBlock → ChainConsensus quorum → Cold tier, it MUST be raw + deterministic. If consumed locally (emotion projection, shard retrieval, consolidation sleep-cycle), it SHOULD be latent. Bridge functions (raw→latent projection, latent→raw scalar clamp) MUST be zero-alloc, gateable, sync-invariant.
KG triple emission: semantic encounters → KG triple from latent similarity. Physical events → TxDelta with raw values, NOT KG triple. Never substitute latent embedding for raw position in anti-cheat.
Spatial cognition (two-brain model): info brain = real MapPos (synced, ground truth). Think brain = per-NPC SpatialBelief (zone KG triple + stale last_known_pos, fog-of-war gated, NOT synced). Bridge is one-way: real position → belief update only when within visible_radius. Confidence decay: sigmoid(-λ * (current_tick - last_observed_tick)). Two brains MUST exist independently — divergence is emergent, not a bug.
Cross-references (read on demand)
Commercial strategy / moat map: the inline short version above suffices for routing decisions. For Super-GOAT novelty gates or "does this become a pillar?" questions, read_fileriir-ai/.docs/README.md (+ 03_pillars/README.md, 04_supergoat_candidates/README.md). Exhaustive moat analysis at riir-ai/.research/003_Commercial_Open_Source_Strategy_Verdict.md (commercially sensitive — read only when inline short version insufficient).
Research next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
A skill your agent uses whenever the user asks to generate, collect, inspect, or prepare early-experience training data (Implicit World Modeling or Self-Reflection, in the sense of arXiv:2510.08558)…
Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.
Research workflow for distilling ML/AI papers into modelless inference primitives, freeze/thaw runtime patterns, latent-space operations, AND model-based training plans across the multi-repo stack. Research is an agent skill from katopz/katgpt-rs. Research workflow for distilling ML/AI papers into modelless inference primitives, freeze/thaw runtime patterns, latent-space operations, AND model-based training plans across the multi-repo stack.
When should I use Research?
Research fits situations like: reading arxiv papers; deciding which repo a paper belongs in; creating .research/ notes; implementing modelless inference primitives.
How do I install Research in Claude Code?
Run `npx skills add katopz/katgpt-rs --skill research -a claude-code`. Or copy the skill folder (.agents/skills/research in katopz/katgpt-rs) into .claude/skills/research in your project. Claude Code loads it when a task matches its description.
How do I install Research in Codex?
Run `npx skills add katopz/katgpt-rs --skill research -a codex`. Or copy the skill folder (.agents/skills/research in katopz/katgpt-rs) into .agents/skills/research in your project. Codex loads it when a task matches its description.
Can I use Research in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add katopz/katgpt-rs --skill research -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/research, .gemini/skills/research, .github/skills/research and .opencode/skills/research in your project.
What does Research need to run?
Going by SKILL.md and its folder, Research needs the command-line tools its instructions call (git).
Does Research access the network?
SKILL.md names 2 domains. In commands or code: r.jina.ai and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Is Research safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
What licence does Research use?
Research is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Research use?
About 20k tokens (SKILL.md is roughly 80k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
What are the alternatives to Research?
Skills that share tags, products or a category with Research: Early Experience Data (OSU-NLP-Group/EarlyExperience, 102 stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and Sentence-Transformers Training Router (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Research?
katopz (a GitHub user) maintains it in katopz/katgpt-rs, which has 135 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 8, 2026.
Source: katopz/katgpt-rs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.