Agent skill

Recsys Pipeline Architect

by wshobson in wshobson/agents

Design composable recommendation, ranking, and feed pipelines using the six-stage Source→Hydrator→Filter→Scorer→Selector→SideEffect framework popularized by xAI's open-sourced X For You algorithm.

MITAuto-check passedAI & LLM Engineering

Install Recsys Pipeline Architect

skills CLI
$ npx skills add wshobson/agents --skill recsys-pipeline-architect -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents recsys-pipeline-architect --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/machine-learning-ops/skills/recsys-pipeline-architect .claude/skills/recsys-pipeline-architect && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
recsys-pipeline-architect
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
999 words
Files
1
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Design composable recommendation, ranking, and feed pipelines using the six-stage Source→Hydrator→Filter→Scorer→Selector→SideEffect framework popularized by xAI's open-sourced X For You algorithm.

  • Works in 3 steps: Single score vs multi-action prediction → Candidate isolation vs joint scoring → Online vs offline batch
  • Building any system that picks the top K items for a (user
  • SKILL.md covers Overview, When to Use This Skill, The Six-Stage Framework and Workflow When Invoked, plus 5 more sections
  • Calls npx

What it does

Recsys Pipeline Architect is an agent skill from wshobson/agents. Design composable recommendation, ranking, and feed pipelines using the six-stage Source→Hydrator→Filter→Scorer→Selector→SideEffect framework popularized by xAI's open-sourced X For You algorithm. Use when building any system that picks "the top K items for a (user, context)" — content feeds, search ranking, RAG rerankers, task prioritizers, notification triage, ad selection.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Retrieval-augmented generation. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Building any system that picks the top K items for a (user
  • Context) — content feeds
  • Task prioritizers
  • Notification triage

Example prompts

  • “the top K items for a (user, context)”
  • “/recsys-pipeline-architect”

Requirements

  • Python 3
  • Node.js

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Single score vs multi-action prediction
  2. Candidate isolation vs joint scoring
  3. Online vs offline batch

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Recsys Pipeline Architect loads about 2k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 999 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 999 words, ~2,025 tokens.

Download SKILL.mdSave it as .claude/skills/recsys-pipeline-architect/SKILL.md (or your agent's skills folder).
name
recsys-pipeline-architect
description
Design composable recommendation, ranking, and feed pipelines using the six-stage Source→Hydrator→Filter→Scorer→Selector→SideEffect framework popularized by xAI's open-sourced X For You algorithm. Use when building any system that picks "the top K items for a (user, context)" — content feeds, search ranking, RAG rerankers, task prioritizers, notification triage, ad selection.

Recsys Pipeline Architect

A spec-and-scaffold skill for building composable recommendation, ranking, and feed pipelines. Encodes the six-stage pattern popularized by xAI's open-sourced For You algorithm (Apache 2.0) and applies it to any "top K for (user, context)" problem.

Overview

Most "recommendation systems" in production aren't exotic ML — they're pipelines: fetch candidates from one or more sources, enrich them with metadata, drop the ineligible, score the rest, sort and pick the top K, then fire async side effects. The pattern is universal. The scoring function and the items change; the pipeline shape doesn't.

This skill is an independent reimplementation of the pattern (MIT) — no code copied from the original.

When to Use This Skill

  • Building any system that returns "the top K items for a user/context"
  • Designing or refactoring a personalized feed (content, search results, notifications)
  • Wrapping an LLM/ML scorer in proper pipeline plumbing (sources, hydration, filters, side effects)
  • Adding multi-action prediction with tunable weights (instead of a single relevance score)
  • Building a RAG retrieval reranker (cheap retrieval → expensive rerank)
  • Designing a task prioritizer or alert triage system

The Six-Stage Framework

#StageJobParallel?
1SourceFetch candidates from one or more originsYes — multiple sources run in parallel
2HydratorEnrich candidates with metadata needed for filtering and scoringYes — independent hydrators run in parallel
3FilterDrop ineligible candidates (blocked, expired, duplicate, ineligible)Sequential — each filter sees fewer items
4ScorerAssign each surviving candidate one or more scoresSequential — later scorers see earlier scores
5SelectorSort by final score, return top KSingle op
6SideEffectCache, log, emit events, update served-historyAsync — must never block the response
Why this exact order
  • Sources before hydration: know what candidates exist before paying to enrich
  • Hydration before filtering: many filters need metadata the source didn't provide
  • Filtering before scoring: scoring is the expensive stage — drop the ineligible first
  • Scorer chain (not single scorer): real systems compose ML scoring + diversity reranking + business rules
  • Selector after scoring: keeps scoring deterministic and cacheable
  • SideEffects last and async: side effects must never block the user response

Workflow When Invoked

Walk the user through eight steps:

  1. Clarify the use case (one round, three questions only if missing): items being ranked, input context, language/runtime
  2. Identify the candidate sources (usually in-network + out-of-network, but single-source also valid)
  3. List required hydrations — for each filter and scorer, what data does it need that the source didn't provide?
  4. List the filters — cheap before expensive, universal before user-specific (duplicate, self, age, block/mute, previously-served, eligibility)
  5. Design the scorer chain — primary ML/heuristic → combiner (multi-action with weights) → diversity → business rules
  6. Selector — sort descending by final score, take top K (or stratified mix)
  7. SideEffects — cache served IDs, emit impression events, update counters, log analytics; all fire-and-forget
  8. Generate the scaffold in the user's stack

Key Trade-offs to Surface

Never default silently on these — they are product decisions disguised as technical ones.

1. Single score vs multi-action prediction
  • Single score: train one model to predict relevance. To change behavior → retrain.
  • Multi-action: predict P(action) for many actions (P(read), P(like), P(share), P(skip), P(report)), combine with weights at serving time. To change behavior → change weights. No retraining.

The X For You algorithm uses multi-action with both positive and negative weights. Recommend multi-action when the user expects to tune frequently.

2. Candidate isolation vs joint scoring
  • Isolated: each candidate scored independently. Deterministic, cacheable.
  • Joint: candidates attend to each other during scoring (e.g., transformer over the whole batch). More expressive but non-deterministic across batches.

Default to isolation. Joint only when there's a specific reason (e.g., explicit batch-aware diversity).

Show full SKILL.md (405 more words)Show less
3. Online vs offline batch
  • Request-time (online): pipeline runs on each request. Latency budget: 100–300ms.
  • Pre-computed (offline batch): pipeline runs periodically, results cached. Lower latency, lower freshness.
  • Hybrid: candidate retrieval offline, ranking online.

Hard Rules

  1. Do not invent benchmark numbers. "How fast is this?" → "depends on workload, run it yourself."
  2. Attribution discipline. Attribute the pattern as "popularized by xAI's open-sourced For You algorithm" / github.com/xai-org/x-algorithm (Apache 2.0).
  3. No trademark use. Don't name the user's artifact "X-like" or use "For You" branding. Use neutral names: "candidate pipeline", "feed pipeline", "ranking pipeline".
  4. Surface trade-offs. Multi-action vs single, isolation vs joint, online vs offline — never default silently.
  5. The generated scaffold must run. No pseudocode passing as code.
  6. Filter order matters. Cheap before expensive. Universal before user-specific.
  7. Side effects never block. Wrap in fire-and-forget patterns (goroutines / promises without await / asyncio tasks).

Anti-Patterns

  • ❌ Scoring before filtering (wastes compute on candidates that will be dropped)
  • ❌ Synchronous side effects (cache writes / impression emits blocking the response)
  • ❌ A single "relevance" score when the product needs multi-objective tuning
  • ❌ Joint scoring as default (non-deterministic, uncacheable, doesn't compose with reranking)
  • ❌ Pseudocode "for illustration" — the scaffold must actually run

Common Use Cases

Content feed (Strapi v5 plugin, TypeScript)

User has a CMS with 50k articles, wants a personalized "for you" feed. Walk through 8 steps → generate a Strapi plugin scaffold with multi-action scoring, author diversity, standard filters, async side-effect lane.

RAG retrieval reranker (Python async)

User's RAG returns top-50 chunks from a vector DB, wants to rerank with a more expensive scorer and return top-5. Single-source pipeline with a scorer chain (cheap retrieval + expensive rerank).

Task prioritizer (FastAPI service)

User has a queue of incoming task suggestions, wants to rank by "what should this user work on next" considering their past patterns. Items reversed (tasks instead of content), same shape applies.

Notification triage (offline-batch job)

User wants a daily digest that picks the top 10 from the last 24h queue. Offline-batch pipeline. Source = queue, filters = age/dedup/eligibility, scorer = urgency × user-affinity, selector = top 10, side effect = email send (still async).

Upstream

This skill is a single-file adapter for the upstream repository, which ships 5 load-on-demand reference docs and 3 runnable example scaffolds (Strapi v5 / Go / Python — every one green on its test suite, 9/9 tests total).

  • Upstream: https://github.com/mturac/recsys-pipeline-architect
  • Release: v0.1.0 (MIT)
  • References: interfaces in 4 languages (TS/Go/Python/Rust), multi-action scoring, candidate isolation, filter cookbook (12 patterns), scorer cookbook
  • Cross-platform install: npx skills add mturac/recsys-pipeline-architect

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/machine-learning-ops/skills/recsys-pipeline-architect of wshobson/agents.

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Recsys Pipeline Architect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Recsys Pipeline Architect compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Recsys Pipeline Architect this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
Convex Suggestopenclaw/clawhub9.5k1 repos~625Automated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
Embeddings via 9Routerdecolua/9router31k—~604Automated safety check: PassMIT
Tavily Search API Integrationandrewyng/context-hub14k—~1.1kAutomated safety check: PassMIT
AI SDK Developmenttrypostit/trypost6921 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Convex Suggest

    openclaw/clawhub

    Suggest the matching Convex component when the user hand-rolls a pattern it already solves (crons, sharded-counter, rate-limiter, storage, search, presence, workflow, RAG, prosemirror-sync).

    9.5k GitHub starsUsed in 1 repo~625 tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    31k GitHub stars~604 tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Tavily Search API Integration

    andrewyng/context-hub

    Guides building Tavily integrations for web search, URL extraction, site crawling and AI-assisted research in Python or JavaScript agent and RAG projects.

    14k GitHub stars~1.1k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • AI SDK Development

    trypostit/trypost

    TRIGGER when working with ai-sdk which is Laravel official first-party AI SDK.

    692 GitHub starsUsed in 1 repo~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Senior Prompt Engineer

    maslennikov-ig/claude-code-orchestrator-kit

    Provides reference guides and Python scripts for prompt optimization, RAG evaluation, and agent orchestration when building or tuning LLM systems.

    260 GitHub starsUsed in 3 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 13 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Distributed Tracing

    wshobson/agents

    Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.

    40k GitHub starsUsed in 12 repos~527 tokens
    Auto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 6 days ago
    Auto-check passed

Questions about Recsys Pipeline Architect

What does Recsys Pipeline Architect do?

Design composable recommendation, ranking, and feed pipelines using the six-stage Source→Hydrator→Filter→Scorer→Selector→SideEffect framework popularized by xAI's open-sourced X For You algorithm. Recsys Pipeline Architect is an agent skill from wshobson/agents. Design composable recommendation, ranking, and feed pipelines using the six-stage Source→Hydrator→Filter→Scorer→Selector→SideEffect framework popularized by xAI's open-sourced X For You algorithm.

When should I use Recsys Pipeline Architect?

Recsys Pipeline Architect fits situations like: building any system that picks the top K items for a (user; context) — content feeds; task prioritizers; notification triage.

How do I install Recsys Pipeline Architect in Claude Code?

Run `npx skills add wshobson/agents --skill recsys-pipeline-architect -a claude-code`. Or copy the skill folder (plugins/machine-learning-ops/skills/recsys-pipeline-architect in wshobson/agents) into .claude/skills/recsys-pipeline-architect in your project. Claude Code loads it when a task matches its description.

How do I install Recsys Pipeline Architect in Codex?

Run `npx skills add wshobson/agents --skill recsys-pipeline-architect -a codex`. Or copy the skill folder (plugins/machine-learning-ops/skills/recsys-pipeline-architect in wshobson/agents) into .agents/skills/recsys-pipeline-architect in your project. Codex loads it when a task matches its description.

Can I use Recsys Pipeline Architect in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill recsys-pipeline-architect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/recsys-pipeline-architect, .gemini/skills/recsys-pipeline-architect, .github/skills/recsys-pipeline-architect and .opencode/skills/recsys-pipeline-architect in your project.

What does Recsys Pipeline Architect need to run?

Going by SKILL.md and its folder, Recsys Pipeline Architect needs the command-line tools its instructions call (npx). Our summary lists: Python 3; Node.js.

Does Recsys Pipeline Architect access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Recsys Pipeline Architect safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Recsys Pipeline Architect use?

Recsys Pipeline Architect is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Recsys Pipeline Architect use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Recsys Pipeline Architect?

Skills that share tags, products or a category with Recsys Pipeline Architect: Convex Suggest (openclaw/clawhub, 9.5k stars), Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Embeddings via 9Router (decolua/9router, 31k stars) and Tavily Search API Integration (andrewyng/context-hub, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Recsys Pipeline Architect?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.