Install the "review-pr" agent skill from https://github.com/zetaalphavector/RAGElo/tree/master/.agents/skills/review-pr into .claude/skills/review-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-pr", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add zetaalphavector/RAGElo --skill review-pr -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "review-pr" agent skill from https://github.com/zetaalphavector/RAGElo/tree/master/.agents/skills/review-pr into .agents/skills/review-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-pr", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add zetaalphavector/RAGElo --skill review-pr -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "review-pr" agent skill from https://github.com/zetaalphavector/RAGElo/tree/master/.agents/skills/review-pr into .cursor/skills/review-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-pr", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add zetaalphavector/RAGElo --skill review-pr -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "review-pr" agent skill from https://github.com/zetaalphavector/RAGElo/tree/master/.agents/skills/review-pr into .gemini/skills/review-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-pr", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add zetaalphavector/RAGElo --skill review-pr -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "review-pr" agent skill from https://github.com/zetaalphavector/RAGElo/tree/master/.agents/skills/review-pr into .github/skills/review-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-pr", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add zetaalphavector/RAGElo --skill review-pr -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "review-pr" agent skill from https://github.com/zetaalphavector/RAGElo/tree/master/.agents/skills/review-pr into .opencode/skills/review-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-pr", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
review-pr
GitHub stars
133
Token cost
~5k tokens
SKILL.md length
2,045 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0
At a glance
Reviews pull requests for code quality, pattern conformance, architecture alignment, security, and completeness.
Works in 10 steps: Gather PR Context → Verify Implementation Matches PR… → Check for Breaking Changes → …
Asked to review a PR
SKILL.md covers Overview, Output Format, Step 1: Gather PR Context and Step 2: Verify Implementation…, plus 5 more sections
Calls gh, jq and git
What it does
Review PR is an agent skill from zetaalphavector/RAGElo. Reviews pull requests for code quality, pattern conformance, architecture alignment, security, and completeness. Use when asked to review a PR, validate code changes, or check if a PR is ready for merge.
Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Development, covering Pull requests. The repository describes itself as: RAGElo is a set of tools that helps you selecting the best RAG-based LLM agents by using an Elo ranker. The licence is Apache-2.0.
When your agent uses it
Asked to review a PR
Validate code changes
Check if a PR is ready for merge
Example prompts
“Use the review-pr skill to review pull requests for code quality, pattern conformance, architecture alignment, security, and completeness”
“/review-pr”
Requirements
Python 3
Workflow steps
10 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6f23f0f. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
Shell commands in SKILL.md call:
gh
jq
git
pytest
ruff
mypy
pip
From the folder's file list and the shell code blocks in SKILL.md.
Network
No URLs in SKILL.md. Its commands use gh, git and pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
Review PR loads about 5k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 2,045 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~53
When it runs· the whole SKILL.md, loaded when a task matches
~5k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Download SKILL.mdSave it as .claude/skills/review-pr/SKILL.md (or your agent's skills folder).
name
review-pr
description
Reviews pull requests for code quality, pattern conformance, architecture alignment, security, and completeness. Use when asked to review a PR, validate code changes, or check if a PR is ready for merge.
PR Review Skill
Overview
RAGElo is a Python library and CLI for evaluating RAG (Retrieval-Augmented Generation) agents using Elo-based tournament ranking. This skill reviews pull requests to ensure they:
Actually implement what the PR description claims
Follow existing codebase patterns (the Chameleon Principle)
Match the code style of the existing codebase
Align with RAGElo's established architecture (Factory+Registry, async evaluators, Pydantic configs)
Address security concerns properly
Introduce no breaking changes to the public API (or have proper migration paths)
Relate to existing GitHub issues appropriately
The Chameleon Principle
The codebase should feel like it was architected by one mind, not assembled by mercenaries.
Every PR must introduce changes that blend seamlessly with existing patterns. When reviewing:
Find the existing pattern first - Before accepting any change, search for how similar problems are already solved
Reject foreign patterns - If the PR introduces a pattern that doesn't exist elsewhere, flag it ruthlessly
Suggest the existing way - Always recommend the established approach over novel solutions
Be ruthless - Pattern violations are not style nits; they are architectural debt
This principle applies to everything: architecture, code organization, naming conventions, testing approaches, and code style. AI agents are notorious for ignoring existing patterns and producing the same generic slop everywhere. The reviewer must catch this.
Output Format
The agent must produce:
Summary - Brief overview of what the PR does
Implementation Verification - Does the code actually do what the PR claims?
Breaking Changes - Any backward-incompatible changes to the public API and migration paths
Pattern & Style Conformance - Does the PR follow existing codebase patterns and style?
Architecture Alignment - Is the PR consistent with RAGElo's Factory+Registry architecture?
Code Quality & Testing - Test coverage, test quality, code organization
Issue Linkage - Related GitHub issues
Security Concerns - Issues identified against security requirements
Recommendations - Required changes, suggestions, and approval status
Step 1: Gather PR Context
Get the PR description
bash
gh pr view --json title,body,number | jq '.'
The PR description tells you what the author claims the PR does. Your job is to verify this claim against the actual code.
Get the list of commits
bash
git log origin/master..HEAD --oneline
Get the full diff
bash
git diff origin/master..HEAD
Get list of changed files
bash
git diff --name-only origin/master..HEAD
CRITICAL: Do not rely on commit messages or PR descriptions alone. You MUST read the actual code changes to understand what was done. The PR description may be incomplete, misleading, or outright wrong.
Accepts LLMProviderConfig (or subclass) in constructor
Registered via @LLMProviderFactory.register(LLMProviderTypes.NAME)
Supports structured output via response_schema parameter
Data model pattern:
Core types (Query, Document, AgentAnswer, PairwiseGame) used as-is — not subclassed unnecessarily
Result types inherit from EvaluatorResult
LLM answer schemas in ragelo/types/answer_formats.py use Pydantic BaseModel
Extra CSV columns → metadata dict pattern respected
Prompt templating pattern:
Jinja2 Template objects, not raw string formatting
Template variables match established names: {{ query.query }}, {{ document.text }}, {{ answer.text }}, {{ game.agent_a_answer.text }}
Metadata accessed via {{ query.metadata.column_name }}
CLI pattern:
Uses Typer (ragelo/cli/)
CLI parameters dynamically generated from Pydantic config classes
Follows existing subcommand structure
Red flags to ruthlessly reject
🚫 Novel patterns when established ones exist
"Let's use a different approach here" → NO. Use the existing approach.
🚫 Inconsistent naming
user_id in one place, userId in another → NO. Match existing convention (snake_case throughout).
🚫 Bypassing the Factory+Registry pattern
Directly instantiating evaluators instead of using get_retrieval_evaluator() → NO. Use the factory.
🚫 New dependencies for solved problems
"Let's add library X" when existing code solves it → NO. Use existing solution.
🚫 Different error handling
Custom exception hierarchies that don't match existing → NO. Use established patterns.
🚫 f-strings or .format() for LLM prompts
All prompts must use Jinja2 Template objects → NO exceptions.
🚫 Synchronous-only evaluator implementations
All evaluators must implement evaluate_async() → NO synchronous-only variants.
When flagging pattern violations
Always show the existing pattern as evidence:
Pattern Violation: PR uses f-string for prompt formatting
Existing Pattern: Codebase uses Jinja2 Template for all LLM prompts
Evidence: Found in ragelo/evaluators/retrieval_evaluators/reasoner_evaluator.py, ragelo/evaluators/answer_evaluators/pairwise_evaluator.py
Required Fix: Convert to Jinja2 Template
Step 6: Code Style Conformance (Detect AI Slop)
AI coding agents are notorious for producing generic, pattern-ignorant code. The codebase must read as if a single person wrote it. Flag these common AI slop indicators:
Linting & Formatting
RAGElo uses ruff for linting and formatting:
Rules: E, F, I (errors, pyflakes, isort)
Line length: 119 characters
Type checking: mypy with pydantic.mypy plugin
Verify with:
bash
ruff check ragelo/
ruff format --check ragelo/
mypy ragelo/
Class Design
Single Responsibility Principle:
Each class has one clear responsibility
Classes are not "god objects" doing everything
Dependency Injection:
BaseLLMProvider is injected into evaluators, not instantiated inside them
Evaluators receive config objects, not raw kwargs
Imports
Absolute imports from ragelo package:
✅ from ragelo.types.evaluables import Query
❌ from .evaluables import Query
❌ from ..types.evaluables import Query
Note:from __future__ import annotations is used throughout the codebase for forward references.
Show full SKILL.md (805 more words)Show less
Comments and Docstrings
Minimal, meaningful documentation:
No useless comments stating the obvious
Docstrings only when function signature is not self-explanatory
Style Violation: Excessive comments/docstrings
python
def get_evaluator(name: str) -> BaseEvaluator:
"""Get an evaluator by name.
Args:
name: The name of the evaluator
Returns:
The evaluator instance
"""
Required Fix: Remove docstring - the function signature is self-explanatory
Async Conventions
Core evaluation logic is in evaluate_async() methods
call_async_fn() utility used for sync→async bridge (from ragelo.utils)
asyncio.wait() used for bounded concurrent execution in evaluate_experiment()
No mixing of sync and async patterns within the same flow
Step 7: Testing Quality
Testing Patterns in RAGElo
RAGElo uses pytest with pytest-asyncio and pytest-mock.
Key testing conventions:
MockLLMProvider in tests/conftest.py returns deterministic responses based on the requested Pydantic schema
experiment fixture loads test data from tests/data/ (2 queries, 4 docs, 2 agents)
OpenAI integration tests gated with @pytest.mark.requires_openai (skipped unless --runopenai)
Tests organized under tests/unit/ and tests/cli/
Test Quality Checklist
Uses MockLLMProvider or similar fixture — not hand-rolled mocks that mirror implementation
Tests the public interface (factory functions, evaluate_experiment()) not internal methods
Core evaluables in ragelo/types/evaluables.py: Query, Document, AgentAnswer, PairwiseGame
Results in ragelo/types/results.py: EvaluatorResult and subclasses
LLM answer schemas in ragelo/types/answer_formats.py: Pydantic models for structured LLM output
Configs in ragelo/types/configurations/: one config class per component
Experiment Orchestration:
Experiment class in ragelo/types/experiment.py is the central orchestrator
Loads from CSV, manages evaluation state, caches to avoid redundant LLM calls, persists to JSON/JSONL
If architecture violations are found:
⚠️ Architecture Violation: [Describe the violation]
Expected pattern: [Explain the correct RAGElo pattern with file references]
Required Fix: [Specific remediation]
Step 9: Security Review
Review security concerns relevant to a library that makes LLM API calls. Do NOT output a checklist. Instead, for each concern identified, use this format:
Security Requirements to Verify
No hardcoded API keys, secrets, or credentials (use environment variables)
API keys handled via SecretStr or similar — not logged or printed
Jinja2 templates not vulnerable to template injection from user-supplied metadata
No arbitrary code execution from user-supplied prompts or configuration
Dependencies properly declared in pyproject.toml with version constraints
Cached results (JSON/JSONL files) don't leak sensitive information
Output Format for Security Issues
For each security concern:
🔴 Security Concern: [Brief title]
Issue: [What the code does wrong or fails to do]
Risk: [Potential exploit, impact, or vulnerability]
Required Fix: [Specific remediation steps]
If no security concerns are identified, state: "No security concerns identified in this review."
Step 10: Generate Review Summary
Compile findings into a structured review:
markdown
# PR Review: [PR Title]
## Summary
[1-2 sentence overview of what the PR does]
## Implementation Verification
[✅ Matches PR description / ⚠️ Discrepancies found]
[List any mismatches between description and code]
## Breaking Changes
[🚨 Breaking changes found / ✅ No breaking changes]
[List any breaking changes with migration path assessment]
## Pattern & Style Conformance
[✅ Passes / ⚠️ Issues Found]
[List any violations with required fixes]
## Architecture Alignment
[✅ Aligned / ⚠️ Violations Found]
[List any architecture violations]
## Code Quality & Testing
[✅ Good / ⚠️ Issues Found]
[List any testing or code quality issues]
## Issue Linkage
- **Related Issues**: [List or "None identified"]
## Security Concerns
[List concerns in Security Concern format, or "No security concerns identified"]
## Recommendations
### Required Changes (Blocking)
1. [Critical issues that must be fixed]
### Suggested Improvements (Non-blocking)
1. [Nice-to-have improvements]
## Approval Status
[✅ Approved / ⚠️ Approved with comments / 🚫 Changes requested]
What NOT to Do
❌ Do not approve without reading the actual code changes
❌ Do not trust PR descriptions without verification
❌ Do not skip pattern and style conformance checking
❌ Do not accept novel patterns when established ones exist
❌ Do not treat pattern violations as style preferences
❌ Do not ignore security concerns (especially API key handling)
❌ Do not allow breaking changes to the public API without migration paths
❌ Do not let the codebase feel like it was built by mercenaries
❌ Do not accept AI slop (generic, pattern-ignorant code)
❌ Do not accept tests that don't use existing fixtures from conftest.py
❌ Do not accept f-strings or .format() for LLM prompt construction
❌ Do not accept synchronous-only evaluator implementations
❌ Do not accept components that bypass the Factory+Registry pattern
What TO Do
✅ Verify code actually implements what the PR description claims
✅ Flag all breaking changes to the public Python API and CLI
✅ Search for existing patterns BEFORE evaluating changes - This is mandatory
✅ Cite specific file:line references when flagging violations
✅ Verify code style matches existing codebase (ruff rules, 119 char lines, snake_case)
✅ Verify new components follow Factory+Registry pattern with enum types
✅ Verify evaluators are async-first and use Jinja2 templates
✅ Verify testing follows existing patterns (MockLLMProvider, fixtures, @pytest.mark.requires_openai)
✅ Check for related GitHub issues
✅ Review API key handling and credential security
✅ Ruthlessly reject foreign patterns - Suggest the existing way instead
✅ Be a chameleon - Changes must blend seamlessly with existing code
✅ Flag monolith page files and demand decomposition into _lib/ and _components/
✅ Flag duplicated UI patterns and utility functions — demand shared components/modules
✅ Flag complex inline JSX callbacks — demand named handler functions
✅ Provide specific, actionable feedback
✅ Distinguish between blocking and non-blocking issues
Review PR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.
Checks a GitHub, GitLab, or Perforce (p4) pull request (or merge request, or shelved changelist) for unresolved review comments, failing status checks, and incomplete PR descriptions.
For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…
Reviews WooCommerce code changes against the project's standards, flagging backend PHP architecture, naming, documentation, data integrity and testing violations.
Reviews pull requests for code quality, pattern conformance, architecture alignment, security, and completeness. Review PR is an agent skill from zetaalphavector/RAGElo. Reviews pull requests for code quality, pattern conformance, architecture alignment, security, and completeness.
When should I use Review PR?
Review PR fits situations like: asked to review a PR; validate code changes; check if a PR is ready for merge.
How do I install Review PR in Claude Code?
Run `npx skills add zetaalphavector/RAGElo --skill review-pr -a claude-code`. Or copy the skill folder (.agents/skills/review-pr in zetaalphavector/RAGElo) into .claude/skills/review-pr in your project. Claude Code loads it when a task matches its description.
How do I install Review PR in Codex?
Run `npx skills add zetaalphavector/RAGElo --skill review-pr -a codex`. Or copy the skill folder (.agents/skills/review-pr in zetaalphavector/RAGElo) into .agents/skills/review-pr in your project. Codex loads it when a task matches its description.
Can I use Review PR in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zetaalphavector/RAGElo --skill review-pr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-pr, .gemini/skills/review-pr, .github/skills/review-pr and .opencode/skills/review-pr in your project.
What does Review PR need to run?
Going by SKILL.md and its folder, Review PR needs the command-line tools its instructions call (gh, jq, git, pytest, ruff and mypy). Our summary lists: Python 3.
Does Review PR access the network?
SKILL.md contains no URLs. Its commands use gh, git and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Is Review PR safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
What licence does Review PR use?
Review PR is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Review PR use?
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
What are the alternatives to Review PR?
Skills that share tags, products or a category with Review PR: Finishing a Development Branch (obra/superpowers, 297k stars), PR Babysitter (openinterpreter/openinterpreter, 69k stars), Check PR (onyx-dot-app/onyx, 32k stars) and PR Design Doc (OpenHands/OpenHands, 91k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Review PR?
zetaalphavector (a GitHub organization) maintains it in zetaalphavector/RAGElo, which has 133 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 24, 2026.
Source: zetaalphavector/RAGElo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.