Agent skill

Streaming Architecture

by majiayu000 in majiayu000/litellm-rs

LiteLLM-RS Streaming Architecture. An agent skill from majiayu000/litellm-rs.

MITAuto-check passedAI & LLM Engineering

Install Streaming Architecture

skills CLI
$ npx skills add majiayu000/litellm-rs --skill streaming-architecture -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/litellm-rs streaming-architecture --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/litellm-rs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/streaming-architecture .claude/skills/streaming-architecture && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
streaming-architecture
GitHub stars
118
Token cost
~2.1k tokens
SKILL.md length
436 words
Files
6
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

LiteLLM-RS Streaming Architecture. An agent skill from majiayu000/litellm-rs.

  • Debugging SSE parsing
  • SKILL.md covers Overview, Core Components and References
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Modifying a provider stream transformer

What it does

Streaming Architecture is an agent skill from majiayu000/litellm-rs. LiteLLM-RS Streaming Architecture. Covers UnifiedSSEParser line buffering, the SSETransformer trait, UnifiedSSEStream backpressure and overflow guarding, provider-specific transformers, and server-side SSE emission. Use when debugging SSE parsing, writing or modifying a provider stream transformer, wiring the stream processing pipeline, or tuning the stream idle timeout.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `reference/best-practices.md`, `reference/buffer-management.md` and `reference/configuration.md`).

It sits in AI & LLM Engineering, covering Model routing and gateways and Backend development. It works with OpenAI and Rust. The repository describes itself as: Self-hosted Rust LLM gateway with OpenAI-compatible APIs, load balancing, failover, and a reusable Rust kernel. The licence is MIT.

When your agent uses it

  • Debugging SSE parsing
  • Modifying a provider stream transformer
  • Wiring the stream processing pipeline
  • Tuning the stream idle timeout

Example prompts

  • “/streaming-architecture”

What it can do on your machine

Read from SKILL.md and the folder at commit ed3f4d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are rust).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Streaming Architecture loads about 2.1k tokens when it runs. Until then it costs about 99 tokens; SKILL.md has 436 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/litellm-rs at commit ed3f4d9, republished under its MIT licence (© majiayu000). 436 words, ~2,113 tokens.

Download SKILL.mdSave it as .claude/skills/streaming-architecture/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
streaming-architecture
description
LiteLLM-RS Streaming Architecture. Covers UnifiedSSEParser line buffering, the SSETransformer trait, UnifiedSSEStream backpressure and overflow guarding, provider-specific transformers, and server-side SSE emission. Use when debugging SSE parsing, writing or modifying a provider stream transformer, wiring the stream processing pipeline, or tuning the stream idle timeout.

Streaming Architecture Guide

Overview

Provider streaming lives in src/core/providers/base/sse.rs plus per-provider transformers under src/core/providers/base/sse/ (openai.rs, anthropic.rs, gemini.rs, cohere.rs, databricks.rs). The layer consumes a provider's raw SSE byte stream and yields Result<ChatChunk, ProviderError> items in an OpenAI-compatible shape, so the server routes never see provider-specific formats.

Streaming Flow
┌─────────────────────────────────────────────────────────────────┐
│                  Provider SSE byte stream                       │
│  reqwest::Response::bytes_stream()                              │
│  (OpenAI, Anthropic, Google, ...)                               │
└─────────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│                  UnifiedSSEStream<S, T>                         │
│  - polls upstream bytes, feeds UnifiedSSEParser                 │
│  - chunk_buffer: VecDeque<ChatChunk>, capped at 10_000          │
│  - Item = Result<ChatChunk, ProviderError>                      │
└─────────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│                  UnifiedSSEParser<T>                            │
│  - String line buffer (incomplete tail retained across reads)   │
│  - SSEEvent field parsing, multi-line data joining              │
│  - end-marker / finish_stream dispatch                          │
└─────────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│                  SSETransformer (per provider)                  │
│  - transform_chunk / transform_stream_chunk                     │
│  - normalizes wire format to ChatChunk                          │
└─────────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│                  Server route re-serialization                  │
│  ChatChunk -> SSE frames ("data: {...}\n\n") + final [DONE]     │
└─────────────────────────────────────────────────────────────────┘

The parser owns its transformer: UnifiedSSEParser<T: SSETransformer> calls back into T while parsing, so there is no separate processing stage between parser and transformer.


Core Components

SSEEvent
rust
// src/core/providers/base/sse.rs
#[derive(Debug, Clone)]
pub struct SSEEvent {
    pub event_type: Option<String>,
    pub data: String,
    pub id: Option<String>,
    pub retry: Option<u64>,
}

SSEEvent::from_line(&str) -> Option<SSEEvent> parses one SSE field line:

  • Empty lines and : comment lines return None.
  • data, event, id, and retry set the matching field; whitespace after the colon is trimmed.
  • retry must parse as u64, otherwise None; unknown fields return None.

The parser accumulates multiple data lines of one event, joining them with \n, and dispatches on the blank line that terminates the event.

SSETransformer Trait
rust
// src/core/providers/base/sse.rs
pub trait SSETransformer: Send + Sync {
    fn provider_name(&self) -> &'static str;

    fn is_end_marker(&self, data: &str) -> bool {
        data.trim() == "[DONE]"
    }

    fn transform_chunk(&self, data: &str) -> Result<Option<ChatChunk>, ProviderError>;

    fn transform_stream_chunk(&self, data: &str) -> Result<Option<ChatChunk>, ProviderError> {
        self.transform_chunk(data)
    }

    fn finish_stream(&self) -> Result<Option<ChatChunk>, ProviderError> {
        Ok(None)
    }

    fn parse_finish_reason(&self, reason: &str) -> Option<FinishReason> { ... }
}
  • Errors are ProviderError (crate::core::providers::unified_provider::ProviderError). There is no dedicated StreamError enum.
  • The default parse_finish_reason maps case-insensitively: stop|end_turn -> Stop, length|max_tokens -> Length, tool_calls|function_call|tool_use -> ToolCalls, content_filter|safety|recitation -> ContentFilter, stop_sequence -> StopSequence, refusal -> Refusal, pause_turn -> PauseTurn; unknown strings yield None.
  • Built-in implementations: OpenAICompatibleTransformer, AnthropicTransformer, GeminiTransformer, CohereTransformer, DatabricksTransformer (see reference/provider-transformers.md).
UnifiedSSEParser<T>
rust
// src/core/providers/base/sse.rs
pub struct UnifiedSSEParser<T: SSETransformer> {
    transformer: T,
    buffer: String,
    current_event: Option<SSEEvent>,
}

impl<T: SSETransformer> UnifiedSSEParser<T> {
    pub fn new(transformer: T) -> Self;
    pub fn process_bytes(&mut self, bytes: &[u8]) -> Result<Vec<ChatChunk>, ProviderError>;
}
  • The buffer is a String, not a byte deque. Each incoming read is decoded independently with String::from_utf8_lossy and appended; only text up to the last \n is processed and the incomplete tail stays buffered for the next call. Line/event splits are retained, but a read boundary inside a multibyte UTF-8 code point is lossy because the undecoded bytes are not retained.
  • process_bytes runs non-stream mode: an end marker produces nothing and events go through transform_chunk.
  • UnifiedSSEStream drives the private process_stream_bytes path (stream mode): an end marker triggers transformer.finish_stream() instead, and data goes through transform_stream_chunk.
  • No size cap applies to this buffer.
  • The private finish_stream flushes any leftover partial line and pending event, then appends transformer.finish_stream() output.
Show full SKILL.md (138 more words)Show less
UnifiedSSEStream<S, T>
rust
// src/core/providers/base/sse.rs
const MAX_CHUNK_BUFFER_SIZE: usize = 10_000;

pub struct UnifiedSSEStream<S, T>
where
    S: Stream<Item = Result<Bytes, reqwest::Error>> + Send + Unpin,
    T: SSETransformer + Clone,
{
    inner: S,
    parser: UnifiedSSEParser<T>,
    chunk_buffer: VecDeque<ChatChunk>,
    pending_error: Option<ProviderError>,
    finished: bool,
}

poll_next order: pop chunk_buffer, then take pending_error, then return None once finished, otherwise poll inner and feed bytes through process_stream_bytes.

  • A read that yields zero complete chunks stores nothing; the stream returns Pending after cx.waker().wake_by_ref().
  • If buffered plus new chunks would exceed MAX_CHUNK_BUFFER_SIZE (10_000), it yields Err(ProviderError::network(...)) instead of growing unboundedly.
  • Transport errors are wrapped as ProviderError::network(provider, format!("Stream error: {error}")); chunks drained from parser.finish_stream() are emitted before the error item.
  • Upstream end-of-stream sets finished and drains parser.finish_stream() before returning None.

Helper create_provider_sse_stream(response, provider_name) boxes response.bytes_stream() behind an OpenAICompatibleTransformer.


References

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in .claude/skills/streaming-architecture of majiayu000/litellm-rs.

  • SKILL.md
  • reference/best-practices.md
  • reference/buffer-management.md
  • reference/configuration.md
  • reference/provider-transformers.md
  • reference/stream-pipeline.md

Open the folder on GitHubat commit ed3f4d9

Compare with similar skills

Streaming Architecture next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Streaming Architecture compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Streaming Architecture this skillmajiayu000/litellm-rs118—~2.1kAutomated safety check: PassMIT
Evaluating Bitrouter Routesbitrouter/bitrouter235—~1.2kAutomated safety check: PassApache-2.0
Run Bitrouter Benchmarkbitrouter/bitrouter235—~2.2kAutomated safety check: PassApache-2.0
Page AgentTommy-yw/RunbookHermes5463 repos~2.3kAutomated safety check: NotesMIT
Embeddings via 9Routerdecolua/9router31k—~604Automated safety check: PassMIT
Reachai Onboardingw8123/EnterpriseAgentFramework865—~6.1kAutomated safety check: PassMIT

Similar skills

  • Evaluating Bitrouter Routes

    bitrouter/bitrouter

    A skill your agent uses when evaluating BitRouter route decisions or Eval Exchange subjects with task-native verifiers, human reviewers, private enterprise evaluators, agentic judges, or genuinely…

    235 GitHub stars~1.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Run Bitrouter Benchmark

    bitrouter/bitrouter

    A skill your agent uses when a user wants to run, compare, resume, audit, share, or submit a Harbor benchmark through BitRouter, including choosing a Harbor dataset and agent, confirming routed…

    235 GitHub stars~2.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Page Agent

    Tommy-yw/RunbookHermes

    Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with…

    546 GitHub starsUsed in 3 repos~2.3k tokens
    Productivity & AutomationAuto-check: notes
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    31k GitHub stars~604 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Reachai Onboarding

    w8123/EnterpriseAgentFramework

    Integrate Java business systems with ReachAI SDK registration, SDK instance heartbeat, gateway/embed access, and optional API Management handoff.

    865 GitHub stars~6.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy API

    starbaser/ccproxy

    Guides users through ccproxy as an OpenAI-compatible and Anthropic-compatible LLM API server with SDK integration, OAuth authentication, sentinel key substitution, model routing, and troubleshooting.

    350 GitHub stars~4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from majiayu000/litellm-rs

All 9 skills in this repo
  • Auth Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Authentication Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Caching Architecture

    majiayu000/litellm-rs

    LiteLLM-RS response caching architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Config Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Configuration Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Error Handling

    majiayu000/litellm-rs

    LiteLLM-RS Error Handling Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Observability Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Provider Architecture

    majiayu000/litellm-rs

    LiteLLM-RS provider system in two tiers - data-driven OpenAI-compatible catalog entries auto-routed through OpenAILikeProvider, plus code-based provider modules implementing the LLMProvider trait…

    118 GitHub stars~4.9k tokensUpdated today
    Auto-check passed

Works with

Questions about Streaming Architecture

What does Streaming Architecture do?

LiteLLM-RS Streaming Architecture. An agent skill from majiayu000/litellm-rs. Streaming Architecture is an agent skill from majiayu000/litellm-rs. LiteLLM-RS Streaming Architecture.

When should I use Streaming Architecture?

Streaming Architecture fits situations like: debugging SSE parsing; modifying a provider stream transformer; wiring the stream processing pipeline; tuning the stream idle timeout.

How do I install Streaming Architecture in Claude Code?

Run `npx skills add majiayu000/litellm-rs --skill streaming-architecture -a claude-code`. Or copy the skill folder (.claude/skills/streaming-architecture in majiayu000/litellm-rs) into .claude/skills/streaming-architecture in your project. Claude Code loads it when a task matches its description.

How do I install Streaming Architecture in Codex?

Run `npx skills add majiayu000/litellm-rs --skill streaming-architecture -a codex`. Or copy the skill folder (.claude/skills/streaming-architecture in majiayu000/litellm-rs) into .agents/skills/streaming-architecture in your project. Codex loads it when a task matches its description.

Can I use Streaming Architecture in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/litellm-rs --skill streaming-architecture -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/streaming-architecture, .gemini/skills/streaming-architecture, .github/skills/streaming-architecture and .opencode/skills/streaming-architecture in your project.

What does Streaming Architecture need to run?

SKILL.md names no scripts, command-line tools or credentials: Streaming Architecture is instructions for the agent only.

Does Streaming Architecture access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Streaming Architecture safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Streaming Architecture use?

Streaming Architecture is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Streaming Architecture use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Streaming Architecture?

Skills that share tags, products or a category with Streaming Architecture: Evaluating Bitrouter Routes (bitrouter/bitrouter, 235 stars), Run Bitrouter Benchmark (bitrouter/bitrouter, 235 stars), Page Agent (Tommy-yw/RunbookHermes, 546 stars) and Embeddings via 9Router (decolua/9router, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Streaming Architecture?

majiayu000 (a GitHub user) maintains it in majiayu000/litellm-rs, which has 118 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 11, 2026.

Source: majiayu000/litellm-rs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.