Ruflo Multi-Agent Orchestration
ruvnet/ruflo
Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.
This skill should be used when the user asks to "share memory between agents", "KV cache compaction for multi-agent", "orchestrator worker context", "latent briefing", "reduce worker tokens"…
$ npx skills add guanyang/open-agent-hub --skill latent-briefing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install guanyang/open-agent-hub latent-briefing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/guanyang/open-agent-hub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/latent-briefing .claude/skills/latent-briefing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "latent-briefing" agent skill from https://github.com/guanyang/open-agent-hub/tree/main/skills/latent-briefing into .claude/skills/latent-briefing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "latent-briefing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/guanyang/open-agent-hub/tree/main/skills/latent-briefingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add guanyang/open-agent-hub --skill latent-briefing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install guanyang/open-agent-hub latent-briefing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guanyang/open-agent-hub.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/latent-briefing .agents/skills/latent-briefing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "latent-briefing" agent skill from https://github.com/guanyang/open-agent-hub/tree/main/skills/latent-briefing into .agents/skills/latent-briefing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "latent-briefing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add guanyang/open-agent-hub --skill latent-briefing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install guanyang/open-agent-hub latent-briefing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guanyang/open-agent-hub.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/latent-briefing .cursor/skills/latent-briefing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "latent-briefing" agent skill from https://github.com/guanyang/open-agent-hub/tree/main/skills/latent-briefing into .cursor/skills/latent-briefing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "latent-briefing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/guanyang/open-agent-hub.git --path skills/latent-briefing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add guanyang/open-agent-hub --skill latent-briefing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install guanyang/open-agent-hub latent-briefing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guanyang/open-agent-hub.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/latent-briefing .gemini/skills/latent-briefing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "latent-briefing" agent skill from https://github.com/guanyang/open-agent-hub/tree/main/skills/latent-briefing into .gemini/skills/latent-briefing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "latent-briefing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install guanyang/open-agent-hub latent-briefingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add guanyang/open-agent-hub --skill latent-briefing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/guanyang/open-agent-hub.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/latent-briefing .github/skills/latent-briefing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "latent-briefing" agent skill from https://github.com/guanyang/open-agent-hub/tree/main/skills/latent-briefing into .github/skills/latent-briefing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "latent-briefing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add guanyang/open-agent-hub --skill latent-briefing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install guanyang/open-agent-hub latent-briefing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guanyang/open-agent-hub.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/latent-briefing .opencode/skills/latent-briefing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "latent-briefing" agent skill from https://github.com/guanyang/open-agent-hub/tree/main/skills/latent-briefing into .opencode/skills/latent-briefing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "latent-briefing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
latent-briefingThis skill should be used when the user asks to "share memory between agents", "KV cache compaction for multi-agent", "orchestrator worker context", "latent briefing", "reduce worker tokens"…
Latent Briefing is an agent skill from guanyang/open-agent-hub. This skill should be used when the user asks to "share memory between agents", "KV cache compaction for multi-agent", "orchestrator worker context", "latent briefing", "reduce worker tokens", "cross-agent memory without summarization", or discusses Attention Matching compaction, recursive language models with workers, or token explosion in hierarchical agents.
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/attention-matching-formulation.md`).
It sits in Agent Workflows, covering Summarization, Agent memory and Multi-agent orchestration. The repository describes itself as: A lightweight, zero-dependency CLI tool to manage and activate capabilities for AI coding assistants (such as Claude Code, Cursor, Trae, etc.). The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit e6ade24. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
arxiv.orgx.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Latent Briefing loads about 3.3k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 95 tokens; SKILL.md has 1,613 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from guanyang/open-agent-hub at commit e6ade24, republished under its MIT licence (© guanyang). 1,613 words, ~3,258 tokens.
.claude/skills/latent-briefing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Hierarchical multi-agent systems often pay for the same context twice. The orchestrator accumulates a long reasoning trajectory, but each worker usually receives only a narrow text handoff such as a subtask prompt plus raw document slices. Passing the full trajectory fixes coverage but drives token cost up on every worker call. Summarization introduces latency and information loss. Retrieval helps with document access but does not preserve the orchestrator's evolving reasoning state.
Latent Briefing addresses this by sharing memory at the representation level rather than the text level. The core idea is to compact the orchestrator trajectory in the worker model's KV cache, keeping positions that are most relevant to the current worker task. The method builds on Attention Matching (AM) KV cache compaction and adapts it for inference-time multi-agent handoff with task-guided queries, a shared token mask across heads, and robust thresholding.
Activate this skill when:
Do not activate this skill for adjacent work owned by other skills:
context-compression, memory-systems, or multi-agent-patterns.memory-systems.multi-agent-patterns.context-optimization.The token explosion pattern. In recursive or REPL-style systems, the orchestrator repeatedly calls a worker to inspect evidence, verify hypotheses, or answer subquestions. The orchestrator's trajectory grows with partial conclusions, dead ends, tool output, and prior worker responses. If that trajectory is passed in full on every worker call, cost compounds quickly.
Representation-level sharing. Instead of summarizing the trajectory into natural language, the system operates on the worker model's KV cache. It retains the positions that the worker would attend to for the current task and drops the rest. This is more specific than ordinary prefix caching: prefix caching reuses identical prefixes, while Latent Briefing also performs task-conditioned selective retention inside the reused trajectory.
Attention Matching as the compaction engine. AM seeks a smaller cache whose attention outputs approximate the full cache. Latent Briefing adapts AM for multi-agent inference by changing the scoring signal and batching strategy:
median + tau * MAD rather than fixed top-k per head.Reference result shape. The public write-up reports substantial worker-token reduction, material total-token savings, and low-single-digit-second compaction overhead on long-document QA workloads (claim-latent-briefing-public-results). Treat these numbers as workload-specific evidence, not a general guarantee.
| Approach | Primary weakness |
|---|---|
| LLM summarization | High latency, lossy abstraction, and no guarantee the summary preserves what the next subtask needs |
| Retrieval / RAG | Depends on chunking and embeddings; can miss cross-chunk or cross-step dependencies |
| Pass full trajectory | Cost scales with every worker call and irrelevant context can degrade worker quality |
Latent Briefing is useful when the bottleneck is not document retrieval itself, but how to transfer orchestrator state into a worker efficiently and precisely.
Frameworks such as Recursive Language Models treat long context as an environment and recurse over it: an orchestrator decomposes work and delegates to workers. Latent Briefing fits the gap where the orchestrator has already built task-specific state that should inform the worker, but re-serializing that state as text is too expensive or noisy.
In the ideal setup, the worker maintains a persistent KV state for the orchestrator trajectory. New trajectory tokens extend that state, then compaction runs just before generation for the current subtask.
Task-guided query vectors. Use queries from the current worker task prompt, not generic samples from the context. Forward-pass the trajectory plus current task through the worker model, then score trajectory positions by how strongly the task attends to them.
Shared token selection. Aggregate scores across layers and heads into one per-position score. One shared mask enables batched operations and avoids hundreds of incompatible per-head solves.
MAD thresholding. Keep positions above a robust outlier threshold such as median + tau * MAD. Higher tau is more aggressive. Optimal settings depend on task regime, trajectory quality, and document length.
Latent Briefing is only practical when the system controls the worker inference runtime closely enough to inspect or transform KV state. It is a poor default for API-only stacks where internal KV tensors are inaccessible. It also assumes the orchestrator trajectory can be represented in the worker's model space. If orchestrator and worker differ materially in tokenizer, architecture, or attention layout, direct representation sharing may not be viable.
Choose the mechanism that matches the bottleneck:
| Need | Prefer | Why |
|---|---|---|
| Stable repeated prefix with minimal logic changes | Prefix caching | Cheapest optimization; no information loss |
| Human-readable and auditable cross-step state | Structured notes or summarization | Easy to inspect and store |
| Sparse lookup across a large external corpus | Retrieval / RAG | Finds documents efficiently |
| Worker needs task-specific slices of orchestrator state and runtime access exists | Latent Briefing | Transfers relevant latent state without replaying all text |
Latent Briefing is not a universal replacement for summarization or retrieval. It is a specialized optimization for systems that already run a controllable orchestrator-worker stack.
Reported long-document QA results suggest:
These are tuning hypotheses, not portable laws. Re-measure on the target workload.
Scenario: orchestrator trajectory grows across worker calls
Call 1: trajectory T1 -> worker answers subquestion A
Call 2: trajectory T2 = T1 + new reasoning + reply A
compact KV(T2) using the task prompt for B
worker answers subquestion BThe task prompt for B decides which parts of T2 survive into the compacted worker state.
Negative example: API-only worker
If the worker runs behind a hosted text-generation API that does not expose KV tensors, Latent Briefing cannot be implemented directly. Use a structured text handoff from context-compression or retrieve state from memory-systems instead.
tau rarely works across long vs short context and easy vs hard tasks. Expect accuracy cliffs when compaction becomes too aggressive.Internal reference:
Related skills in this collection:
External resources:
Created: 2026-04-14 Last Updated: 2026-05-15 Author: Agent Skills for Context Engineering Contributors; primary technical source Ramp Labs (public post) Version: 1.2.0
© guanyang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/latent-briefing of guanyang/open-agent-hub.
Open the folder on GitHubat commit e6ade24
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in guanyang/open-agent-hub, which our catalogue first saw on October 7, 2026.
Latent Briefing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Latent Briefing this skillguanyang/open-agent-hub | 977 | 1 repos | ~3.3k | Automated safety check: Pass | MIT | |
| Ruflo Multi-Agent Orchestrationruvnet/ruflo | 74k | 1 repos | ~975 | Automated safety check: Pass | MIT | |
| Harness Engineering10xChengTu/harness-engineering | 102 | 1 repos | ~1k | Automated safety check: Pass | None | |
| Munder Difflin Hive SyncHarnessMD/munder-difflin | 8.6k | — | ~331 | Automated safety check: Notes | MIT | |
| Session Recaprohitg00/agentmemory | 29k | — | ~510 | Automated safety check: Pass | Apache-2.0 | |
| Marm InitLyellr88/marm-memory | 418 | — | ~7.5k | Automated safety check: Notes | Apache-2.0 |
ruvnet/ruflo
Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.
10xChengTu/harness-engineering
Set up and improve harness engineering (AGENTS.md, docs/, lint rules, eval systems, project-level prompt engineering) for AI-agent-friendly codebases.
HarnessMD/munder-difflin
Runs the start-of-task routine for an agent working in a multi-agent hive: reads memory.md, processes new inbox messages and reminds the agent to record durable facts before ending.
rohitg00/agentmemory
Summarizes recent agent sessions for the current project by date, with each session's title, status, observation count and top highlights from memory.
Lyellr88/marm-memory
Guided MARM MCP setup. An agent skill from Lyellr88/marm-memory.
xvirobotics/metabot
Documents the unified `metabot` CLI for personal memory, the skill hub, durable agent messaging, the agent registry, T5T status and scheduling.
guanyang/open-agent-hub
This skill should be used when long-running agent sessions need context compression, structured summarization, compaction, token-per-task optimization, or durable handoff summaries that preserve…
guanyang/open-agent-hub
This skill should be used to explain or reason about the foundational concepts of context engineering: what context is, the anatomy of a context window, how attention mechanics work, the U-shaped…
guanyang/open-agent-hub
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…
guanyang/open-agent-hub
This skill should be used when designing multi-agent systems that need context isolation, supervisor or swarm coordination, explicit handoffs, parallel execution, or a decision on whether multiple…
guanyang/open-agent-hub
This skill should be used for project-level decisions about LLM-powered systems: whether an LLM is the right primitive for the task at hand, the shape of a multi-stage batch or agent pipeline, token…
guanyang/open-agent-hub
This skill should be used for the tool-interface layer of an agent system specifically: writing tool descriptions agents can route on, designing tool schemas and response formats, naming…
Categories
This skill should be used when the user asks to "share memory between agents", "KV cache compaction for multi-agent", "orchestrator worker context", "latent briefing", "reduce worker tokens"…. Latent Briefing is an agent skill from guanyang/open-agent-hub. This skill should be used when the user asks to "share memory between agents", "KV cache compaction for multi-agent", "orchestrator worker context", "latent briefing", "reduce worker tokens", "cross-agent memory without summarization", or discusses Attention Matching compaction, recursive language models with workers, or token explosion in hierarchical agents.
Latent Briefing fits situations like: asks to share memory between agents; KV cache compaction for multi-agent; orchestrator worker context; latent briefing.
Run `npx skills add guanyang/open-agent-hub --skill latent-briefing -a claude-code`. Or copy the skill folder (skills/latent-briefing in guanyang/open-agent-hub) into .claude/skills/latent-briefing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add guanyang/open-agent-hub --skill latent-briefing -a codex`. Or copy the skill folder (skills/latent-briefing in guanyang/open-agent-hub) into .agents/skills/latent-briefing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guanyang/open-agent-hub --skill latent-briefing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/latent-briefing, .gemini/skills/latent-briefing, .github/skills/latent-briefing and .opencode/skills/latent-briefing in your project.
SKILL.md names no scripts, command-line tools or credentials: Latent Briefing is instructions for the agent only.
SKILL.md names 2 domains. As links in the text: arxiv.org and x.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Latent Briefing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 741 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Latent Briefing: Ruflo Multi-Agent Orchestration (ruvnet/ruflo, 74k stars), Harness Engineering (10xChengTu/harness-engineering, 102 stars), Munder Difflin Hive Sync (HarnessMD/munder-difflin, 8.6k stars) and Session Recap (rohitg00/agentmemory, 29k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
guanyang (a GitHub user) maintains it in guanyang/open-agent-hub, which has 977 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on October 11, 2026.
Source: guanyang/open-agent-hub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.