Agent skill

Rlt Refactor

by ThinkFlowLab in ThinkFlowLab/vllm-rlt

Review and refactor inference-runtime code using concrete rules for responsibility boundaries, state ownership, interfaces, asynchronous lifetimes, KV management, and maintainability.

Apache-2.0Auto-check passedDevelopment

Install Rlt Refactor

skills CLI
$ npx skills add ThinkFlowLab/vllm-rlt --skill rlt-refactor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ThinkFlowLab/vllm-rlt rlt-refactor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ThinkFlowLab/vllm-rlt.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/rlt-refactor .claude/skills/rlt-refactor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rlt-refactor
GitHub stars
149
Token cost
~4.2k tokens
SKILL.md length
2,053 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Review and refactor inference-runtime code using concrete rules for responsibility boundaries, state ownership, interfaces, asynchronous lifetimes, KV management, and maintainability.

  • Code-quality reviews and behavior-preserving refactoring of vllm-rlt
  • SKILL.md covers Responsibilities and State, Interfaces, Types, and…, Boundaries and Capabilities and Lifecycle and Runtime Semantics, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Refactoring

What it does

Rlt Refactor is an agent skill from ThinkFlowLab/vllm-rlt. Review and refactor inference-runtime code using concrete rules for responsibility boundaries, state ownership, interfaces, asynchronous lifetimes, KV management, and maintainability. Use for code-quality reviews and behavior-preserving refactoring of vllm-rlt.

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Refactoring, LLM inference and serving and Code quality. It works with vLLM. The repository describes itself as: vllm based inference engine for looped transformer. The licence is Apache-2.0.

When your agent uses it

  • Code-quality reviews and behavior-preserving refactoring of vllm-rlt
  • Tasks that involve Refactoring
  • Tasks that involve LLM inference and serving

Example prompts

  • “/rlt-refactor”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit b599dc5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rlt Refactor loads about 4.2k tokens when it runs. Until then it costs about 69 tokens; SKILL.md has 2,053 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ThinkFlowLab/vllm-rlt at commit b599dc5, republished under its Apache-2.0 licence (© ThinkFlowLab). 2,053 words, ~4,163 tokens.

Download SKILL.mdSave it as .claude/skills/rlt-refactor/SKILL.md (or your agent's skills folder).
name
rlt-refactor
description
Review and refactor inference-runtime code using concrete rules for responsibility boundaries, state ownership, interfaces, asynchronous lifetimes, KV management, and maintainability. Use for code-quality reviews and behavior-preserving refactoring of vllm-rlt.

Code Quality and Refactoring

Apply the following rules to the changed code and its callers. Explain each applicable finding with its location, triggering condition, impact, and correction. For review-only requests, report findings without modifying code. When implementation is requested, keep changes within scope and validate affected contracts.

Examples illustrate contracts, not mandatory APIs or complete implementations. Introduce a class, enum, helper, or abstraction only when it resolves a concrete problem. Preserve stable rule IDs as this library evolves.

Responsibilities and State

R01 — Separate responsibilities before splitting files

Separate policy, admission, allocation, execution, and result application where they have different responsibilities. Moving code between files without clarifying its dependencies is not enough. Prefer small, cohesive operations over speculative frameworks.

Example: make the scheduling path read as choose_stage() → admit_requests() → build_batch(). Do not hide all three responsibilities inside a renamed SchedulerHelper.run().

R02 — Give mutable state one authoritative owner

Identify who initializes, advances, resets, and releases each state. Distinguish configuration from runtime progress. Other modules should request transitions or provide feedback rather than independently mutate the owner's bookkeeping.

Example: a policy owns its fairness counter and no-refill phase; Scheduler owns queues and logical request progression; ModelRunner owns device tensors and execution events. These owners may use internal helpers without creating additional authoritative copies.

R03 — Match state lifetime to the operation it describes

Keep temporary selections and exclusions local to the operation whose validity they describe. Pass the context explicitly. Check repeated calls, alternate entry points, cancellation, and request-ID reuse; do not depend on a reset performed by only one caller.

Example: a batch builder creates a local selected set and passes excluded=frozenset(selected) when requesting preemption. Storing it on Scheduler can accidentally protect requests selected by an earlier batch.

R04 — Keep policy decisions separate from resource operations

Keep eligibility, ranking, victim selection, and queue transitions with scheduling. Snapshot, copy, restore, and release mechanisms should execute a defined operation and report its outcome. Agree on the transition protocol so failed resource operations do not leave inconsistent scheduling state.

Example: Scheduler chooses a preemption victim and coordinates its suspension. A state-preservation component saves the victim's device state; it does not independently choose a different victim or edit scheduler queues.

Interfaces, Types, and Readability

Name related values crossing module boundaries. Define units, ownership, lifetime, and mutation rights. Specify tensor shape, dtype, device, layout, and borrowing scope where relevant. Do not expose mutable internal domain objects merely for convenience.

Example: AdmissionPlan(required_blocks=8, reclaimable_blocks=3) communicates more than (8, 3). An execution result should identify whether it contains gate scores, sampled tokens, or completed prefill ranges rather than requiring the caller to infer an untyped list's meaning.

R06 — Represent finite states and distinct outcomes explicitly

Use an enum or named outcome when strings or overloaded booleans obscure valid states. Preserve external serialization at the boundary. Ordinary two-way conditions do not automatically need enums.

Example: use ResumeResult.NOT_FOUND, BLOCKED, and RESTORED for three restoration outcomes. An internal FinishReason.LENGTH can still serialize to the existing public string "length".

R07 — Make arguments, units, side effects, and failure outcomes explicit

Expose meaningful dependencies instead of reading unrelated mutable state implicitly. Prefer keyword-only switches where positional booleans obscure meaning. Distinguish valid zero values from failure and define inclusive/exclusive boundaries.

Example:

python
available_count = ensure_active_slot(request, active_count)
if available_count is None:
    return  # Zero is a valid count, not failure.

reserved_blocks = reserved_growth_blocks(reserve_outputs=True)

A frontier for token range [4, 8) is the exclusive token extent 8, not the loop depth. A planning function that updates prefix LRU state must disclose that side effect rather than appear pure.

R08 — Make control flow read as meaningful operations

Extract cohesive steps that reveal intent and keep successful, blocked, retry, and failure paths understandable. Preserve ordering and side effects. Avoid helpers that add indirection without clarifying a responsibility.

Example: admission can read as queue ordering, slot availability, restoration, budget planning, and allocation/queue commit. Extracting those steps must not silently introduce another-stage fallback when the selected stage produces an empty batch.

R09 — Name purpose, scope, and measurement units

Choose names that distinguish quantities and lifecycle roles. Separate supplied input from derived usable extent. Preserve externally defined model fields unless changing them is an intentional compatibility decision.

Example: prefer _pending_exit_signals to _signals, request_id to rid, and watermark_ratio / watermark_blocks to one ambiguous watermark. Keep publishable_tokens separate from the supplied prefilled extent.

R10 — Keep dependencies visible and initialization invariants intact

Initialize required components through their defined lifecycle. Do not weaken production invariants solely to support tests that bypass constructors. Keep ordinary imports visible; use lazy imports when justified by optional dependencies or initialization constraints.

Example: if every initialized engine owns a preemption manager, call self.preemption.discard_snapshot(request_id) directly. A test using object.__new__ should supply that required component rather than require production hasattr guards.

Boundaries and Capabilities

R11 — Replace private-state access with an owner-defined operation

Use narrow public operations or read-only views to cross component boundaries. Define borrowing and mutation rights. A getter exposing unrestricted mutable internals does not resolve ownership. Let the responsible module define its contract rather than creating a competing interface in each caller.

Example: call preemption.discard_snapshot(request_id) instead of preemption.snapshots.pop(request_id, None). Serving should call engine request/statistics methods instead of requiring fake scheduler/cache objects on a PD facade.

R12 — Separate logical KV resources from device operations

Logical KV management owns allocation, references, block mappings, depth validity, leases, and release conditions. Execution owns tensor storage, writes, copies, and buffer/event lifetimes. Attention owns its capability and metadata semantics. Cross-boundary operations need explicit completion feedback.

Example: a scheduler asks for the resource cost of growth without inspecting private reference counts. An early-exit finalization plan describes affected depths and positions; execution performs the copies, and logical validity advances only after the required completion evidence.

R13 — Declare capabilities instead of inferring them from versions

Let each backend or adapter declare the operations it supports. Route through those capabilities. A version identity can inform the adapter's implementation but should not substitute for the caller-facing contract.

Example:

python
# Coupled to one implementation's version numbering.
packed_prefill = backend.generation == 4

# Describes the operation the caller actually requires.
packed_prefill = backend.capabilities.supports_packed_prefill
R14 — Select backends against actual execution requirements

Automatic selection should consider hardware, available implementations, dtype, head dimension, block size, and required execution capabilities. Keep explicit selection available, fail clearly on incompatibility, and report the selected implementation and reason.

Example: an installed package is not enough to establish CUDA Graph or packed-prefill support. An auto request may select a compatible alternative with a diagnostic; an explicit incompatible choice should not silently switch to another backend.

R15 — Share contracts without forcing identical execution timing

Share scheduling, execution, and result-application semantics where they agree. Preserve differences between submission, device completion, and output delivery. Share recurrent computation while keeping prefill/decode state handling explicit at suitable boundaries.

Example: sync and async runners can consume the same task/result contracts while async execution retains multiple in-flight submissions. Requiring both to wait and return at identical points can destroy overlap. A local executor boundary does not, by itself, justify adding generic RPC or unused distributed implementations.

Lifecycle and Runtime Semantics

R16 — Centralize cleanup without losing release ordering

Share cleanup through owner-defined operations across completion, abort, cancellation, and failure. Distinguish immediately releasable state from resources still referenced by asynchronous work. Preserve distinct termination outcomes and required ordering.

Example: abort can delegate common queue/accounting cleanup to finish(reason=ABORT). Recycling a runner slot still waits for the event that covers its last use; consolidating cleanup must not bypass that dependency.

R17 — Reject stale results using execution identity

Associate results with the generation or execution that produced them. Validate that identity before updating current state. An external request ID can be reused and is not sufficient on its own.

Example:

python
if requests.get(result.request_id) is not submitted_request:
    return  # Cancelled request or a different request reusing the ID.

Object identity is appropriate only within a shared process. Cross-process results need an equivalent generation/sequence contract.

Show full SKILL.md (803 more words)Show less
R18 — Distinguish submission, completion, validity, and success

Define the evidence required for each lifecycle transition. A submitted operation is not completed; completed device work does not prove every required range is valid; finished processing does not necessarily mean success.

Examples:

  • Prefix publication needs completed writes and contiguous coverage at every required layer/depth. If only KV is cached, the final prompt token may need recomputation to recover its hidden state.
  • A PD commit marks the end of submission. Destination activation also requires receiving the expected data and satisfying execution dependencies.
  • An artifact job can be complete with an error. Delete recoverable raw files only after successful output verification and publication.
R19 — Preserve loop, exit, restoration, and sampling semantics

Keep token position distinct from loop depth. Preserve gate-score production/consumption timing, cumulative exit state, depth-aware KV validity, and subsequent-token behavior. Restoration must preserve relevant KV, hidden state, logical progress, exit state, and RNG state.

Example: moving sampling into a separate component must preserve request-local generator creation and advancement; resuming a request must not reseed it. Recomputing at full depth is not automatically equivalent to restoring a request that previously exited early.

R20 — Specify handoff and failure protocols across workers

Define valid KV ranges, destination mappings, final prompt hidden state, generation/chunk identities, commit, receive completion, activation, and acknowledgment. Specify how remote writes terminate before memory can be reused after cancellation, timeout, or worker failure.

Preserve interactions with prefix hits, in-flight transfers, and preemption. Do not alter transfer direction or chunk timing merely to fit a new interface.

Example: the P worker sends the required KV and final prompt hidden state; D performs coda and sampling. Cancelling D cannot immediately release destination pages while P may still write to them.

R21 — Preserve device residency and resource/thread ownership

Do not add CPU round trips, .item(), synchronization, or repeated metadata allocation merely to simplify an interface. Keep buffers, streams, events, graphs, and asynchronous state under coherent lifetime contracts. Offload only operations that are safe on the destination thread.

Example: keep sampled token IDs on the GPU when the next prelude consumes them there. A profiler may need finalization/export on its owner thread while parsing and compressing already exported files runs in the background; wrapping everything in a future does not remove that constraint.

Change Discipline and Validation

R22 — Make behavioral fixes explicit within refactoring work

Separate structural changes from intentional behavior changes in the explanation and validation. Describe the old trigger, old outcome, new behavior, and evidence. They may share a change when appropriate, but a behavior fix must not masquerade as a pure extraction.

Example: changing fairness accounting so empty batches no longer consume an allowance is a behavior fix. Specify that the counter tracks nonempty scheduled batches, not GPU completions, and exercise both empty and nonempty cases.

R23 — Migrate incrementally and preserve explicit compatibility contracts

Keep increments runnable. Use temporary adapters only where needed, migrate callers, and define when old paths can be removed. Avoid maintaining two authoritative state stores. Distinguish Python keywords, CLI flags, output serialization, and defaults when assessing compatibility.

Example: renaming a Python keyword from watermark to watermark_ratio requires caller migration even if --kv-watermark stays unchanged. Renaming publish_prefix(rid=...) changes keyword callers even when positional calls still work.

R24 — Explain rationale, contracts, and blocked paths

Document why ordering and lifecycle constraints exist, what state they protect, and what happens when progress is blocked. Use worked examples for budgets/layouts and diagrams where timing is otherwise hard to follow. Keep comments tied to actual invariants rather than narrating obvious syntax.

Update usage documentation when interfaces change. Place detailed evidence where it helps review; do not automatically add a permanent walkthrough or benchmark directory for every extraction.

Example: explain that the final prompt token is excluded from prefix reuse because its hidden state is needed for coda. Show a concrete admission-budget calculation that distinguishes physical blocks, reserved growth, and watermark headroom.

R25 — Validate affected invariants and state the limits of evidence

Validate the changed contract and its callers, including applicable blocked, cancellation, late-result, ID-reuse, failure, and restoration paths. Reuse existing checks where suitable. For a behavioral fix, reproduce the old failure where practical; do not write tests that merely mirror helper structure.

Cover affected exit policies, KV layouts, sampling paths, and feature combinations. Use justified numerical tolerances; do not assume cross-batch bitwise equivalence. For hot-path changes, compare performance under matched conditions and report the measured tradeoff.

Tie evidence to the tested revision and environment. Mocks, CPU tests, skipped GPU tests, and results from an earlier revision do not establish current device behavior or performance. Mark unverified paths explicitly.

Example: test that cancelled work cannot mutate a new request with the same ID, not merely that a helper was called. A CPU test of preemption logic does not establish safe CUDA stream ordering. Documentation-only changes do not require GPU benchmarks.

© ThinkFlowLab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/rlt-refactor of ThinkFlowLab/vllm-rlt.

Open the folder on GitHubat commit b599dc5

Compare with similar skills

Rlt Refactor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rlt Refactor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rlt Refactor this skillThinkFlowLab/vllm-rlt149—~4.2kAutomated safety check: PassApache-2.0
Modern JavaScript Patternswshobson/agents40k12 repos~548Automated safety check: PassMIT
Dynamo Kv Replay Parityai-dynamo/dynamo8.3k—~4.9kAutomated safety check: PassApache-2.0
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Ponytail Lazy Developer ModeDietrichGebert/ponytail160k1 repos~873Automated safety check: PassMIT
Systematic Code Refactoringluongnv89/claude-howto42k—~3kAutomated safety check: PassMIT

Similar skills

  • Covers ES6+ syntax and functional patterns for refactoring older JavaScript: async/await, destructuring, spread, modules, generators and data pipelines.

    40k GitHub starsUsed in 12 repos~548 tokens
    DevelopmentAuto-check passed
  • Dynamo Kv Replay Parity

    ai-dynamo/dynamo

    Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…

    8.3k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Ponytail Lazy Developer Mode

    DietrichGebert/ponytail

    Makes the agent pick the laziest solution that works: skip unneeded work, reuse what exists, prefer the standard library and platform features, and keep diffs small.

    160k GitHub starsUsed in 1 repo~873 tokens
    DevelopmentAuto-check passed
  • Systematic Code Refactoring

    luongnv89/claude-howto

    Guides refactoring in phases based on Martin Fowler's method: research, test coverage check, planning and small tested steps, with your approval at each phase.

    42k GitHub stars~3k tokensUpdated today
    DevelopmentAuto-check passed
  • Dignified Python Standards

    docling-project/docling

    Applies opinionated production Python conventions chosen by the project's Python version: modern type syntax, pathlib, explicit checks and interface guidance.

    69k GitHub stars~1.5k tokensUpdated today
    DevelopmentAuto-check passed

More from ThinkFlowLab/vllm-rlt

  • Vllm Rlt Review PR

    ThinkFlowLab/vllm-rlt

    Review PRs and local changes for hsliuustc0106/vllm-rlt: Ouro engine and KV correctness, serving behavior, and BF16 accuracy/speed evidence.

    149 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rlt Perf Opt

    ThinkFlowLab/vllm-rlt

    Analyze and optimize vllm-rlt inference performance using reproducible unprofiled benchmarks, paired ops-only/full profiles, source-level attribution, and correctness checks.

    149 GitHub stars~4.6k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Rlt Refactor

What does Rlt Refactor do?

Review and refactor inference-runtime code using concrete rules for responsibility boundaries, state ownership, interfaces, asynchronous lifetimes, KV management, and maintainability. Rlt Refactor is an agent skill from ThinkFlowLab/vllm-rlt. Review and refactor inference-runtime code using concrete rules for responsibility boundaries, state ownership, interfaces, asynchronous lifetimes, KV management, and maintainability.

When should I use Rlt Refactor?

Rlt Refactor fits situations like: code-quality reviews and behavior-preserving refactoring of vllm-rlt; tasks that involve Refactoring; tasks that involve LLM inference and serving.

How do I install Rlt Refactor in Claude Code?

Run `npx skills add ThinkFlowLab/vllm-rlt --skill rlt-refactor -a claude-code`. Or copy the skill folder (.agents/skills/rlt-refactor in ThinkFlowLab/vllm-rlt) into .claude/skills/rlt-refactor in your project. Claude Code loads it when a task matches its description.

How do I install Rlt Refactor in Codex?

Run `npx skills add ThinkFlowLab/vllm-rlt --skill rlt-refactor -a codex`. Or copy the skill folder (.agents/skills/rlt-refactor in ThinkFlowLab/vllm-rlt) into .agents/skills/rlt-refactor in your project. Codex loads it when a task matches its description.

Can I use Rlt Refactor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ThinkFlowLab/vllm-rlt --skill rlt-refactor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rlt-refactor, .gemini/skills/rlt-refactor, .github/skills/rlt-refactor and .opencode/skills/rlt-refactor in your project.

What does Rlt Refactor need to run?

SKILL.md names no scripts, command-line tools or credentials: Rlt Refactor is instructions for the agent only. Our summary lists: Python 3.

Does Rlt Refactor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Rlt Refactor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Rlt Refactor use?

Rlt Refactor is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rlt Refactor use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Rlt Refactor?

Skills that share tags, products or a category with Rlt Refactor: Modern JavaScript Patterns (wshobson/agents, 40k stars), Dynamo Kv Replay Parity (ai-dynamo/dynamo, 8.3k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars) and Ponytail Lazy Developer Mode (DietrichGebert/ponytail, 160k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rlt Refactor?

ThinkFlowLab (a GitHub organization) maintains it in ThinkFlowLab/vllm-rlt, which has 149 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 11, 2026.

Source: ThinkFlowLab/vllm-rlt on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.