Eval Guide
microsoft/eval-guide
Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.
Plan and run conversational AI agent evaluations with test generation and analysis.
$ npx skills add mikeyobrien/rho --skill eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mikeyobrien/rho eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mikeyobrien/rho.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval .claude/skills/eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval" agent skill from https://github.com/mikeyobrien/rho/tree/main/skills/eval into .claude/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mikeyobrien/rho/tree/main/skills/evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mikeyobrien/rho --skill eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mikeyobrien/rho eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/rho.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval .agents/skills/eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval" agent skill from https://github.com/mikeyobrien/rho/tree/main/skills/eval into .agents/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mikeyobrien/rho --skill eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mikeyobrien/rho eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/rho.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval .cursor/skills/eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval" agent skill from https://github.com/mikeyobrien/rho/tree/main/skills/eval into .cursor/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mikeyobrien/rho.git --path skills/eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mikeyobrien/rho --skill eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mikeyobrien/rho eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/rho.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval .gemini/skills/eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval" agent skill from https://github.com/mikeyobrien/rho/tree/main/skills/eval into .gemini/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mikeyobrien/rho evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mikeyobrien/rho --skill eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mikeyobrien/rho.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval .github/skills/eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval" agent skill from https://github.com/mikeyobrien/rho/tree/main/skills/eval into .github/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mikeyobrien/rho --skill eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mikeyobrien/rho eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/rho.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval .opencode/skills/eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval" agent skill from https://github.com/mikeyobrien/rho/tree/main/skills/eval into .opencode/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evalPlan and run conversational AI agent evaluations with test generation and analysis.
Eval is an agent skill from mikeyobrien/rho. Plan and run conversational AI agent evaluations with test generation and analysis.
Its SKILL.md is about 9.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering Test generation, Agent evaluation and testing and Chatbots and conversational support. The repository describes itself as: An AI agent that stays running, remembers across sessions, and checks in on its own. macOS, Linux, Android. Built on Pi. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 073a3ee. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvgitpythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval loads about 9.7k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 3,215 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mikeyobrien/rho at commit 073a3ee, republished under its MIT licence (© mikeyobrien). 3,215 words, ~9,664 tokens.
.claude/skills/eval/SKILL.md (or your agent's skills folder).EvalKit is a conversational evaluation framework for AI agents that guides you through creating robust evaluations using the Strands Evals SDK. Through natural conversation, you can plan evaluations, generate test data, execute evaluations, and analyze results.
./chatbot-agent, /path/to/my-agent)Constraints for parameter acquisition:
When a user requests evaluation (any phase), first validate the environment:
Folder Structure:
All evaluation artifacts MUST be created in the eval/ folder at the same level as the target agent folder:
<agent-evaluation-project>/ # Example name - can be any name for user's evaluation project
├── <target-agent-folder>/ # Example name - this is the agent you are evaluating
│ └── [agent source code] # Existing agent code
└── eval/ # All evaluation files go here (sibling to target-agent-folder)
├── eval-plan.md
├── test-cases.jsonl
├── results/
├── run_evaluation.py
├── eval-report.md
└── README.mdNote:
eval/ folder is a sibling directory to user's agent folder, not nested inside itagent-evaluation-project and target-agent-folder are placeholder names - user may use any names that fit their projectConstraints:
When to Trigger: User requests evaluation planning or mentions creating/designing an evaluation
User Intent Recognition:
Execution Flow:
Parse user request: Extract agent path, evaluation focus, and specific requirements from natural language
Navigate to evaluation project directory:
cd <your-evaluation-project> # Navigate to the directory containing both agent folder and eval/Create evaluation directory structure:
mkdir -p evalFollow this execution flow:
Write the complete evaluation plan to eval/eval-plan.md using the template structure (see Appendix A: Evaluation Plan Template), replacing placeholders with concrete details derived from the analysis while preserving section order and headings.
Report completion with evaluation plan file path, and suggest next step: "Would you like me to generate test cases based on this plan?"
High-Level Design (What & Why):
Low-Level Implementation (How):
Evaluation metrics must be:
Key Principles:
eval/ directory structureExamples of reasonable defaults:
Constraints:
When to Trigger: User requests test case generation or mentions creating test data
User Intent Recognition:
Execution Flow:
Parse user request: Extract any specific requirements (e.g., "focus on edge cases", "10 test cases")
Navigate to evaluation project directory:
cd <your-evaluation-project> # Navigate to the directory containing both agent folder and eval/Load the current evaluation plan (eval/eval-plan.md) to understand evaluation areas and test data requirements.
Follow this execution flow:
eval/test-cases.jsonlReport completion with test case count, coverage summary, and suggest next step: "Would you like me to run the evaluation with these test cases?"
Constraints:
When to Trigger: User requests evaluation execution or mentions running tests
User Intent Recognition:
Execution Flow:
Parse user request: Extract any specific requirements (e.g., "run on subset", "verbose output")
Navigate to evaluation project directory:
cd <your-evaluation-project> # Navigate to the directory containing both agent folder and eval/Load the current evaluation plan (eval/eval-plan.md) to understand evaluation requirements and agent architecture.
Follow this execution flow:
requirements.txt at repository root, adding Strands Evals SDK dependenciesuv to create virtual environment, activate it, and install requirements.txteval/run_evaluation.py using Strands Evals SDK patterns with Case objects, Experiment class, and appropriate evaluatorseval/results/ directoryeval/README.md with running instructions for usersReport completion with evaluation results summary and suggest next step: "Would you like me to analyze these results and provide recommendations?"
CRITICAL: Always Create Minimal Working Version: Implement the most basic version that works
CRITICAL REQUIREMENT - Getting Latest Documentation: Before implementing evaluation code, you MUST retrieve the latest Strands Evals SDK documentation and API usage examples. This is NOT optional. You MUST NOT proceed with implementation without either context7 access or the source code. This ensures you're using the most current patterns and avoiding deprecated APIs.
Step 1: Check Context7 MCP Availability: First, check if context7 MCP server is available by attempting to use it. If you receive an error indicating context7 is not available, proceed to Step 3.
Step 2: Primary Method - Using Context7 (If Available):
Step 3: Fallback Method - REQUIRED If Context7 Is Not Available: If context7 MCP is not installed or doesn't have Strands Evals SDK documentation, you MUST STOP and prompt the user to take one of these actions:
REQUIRED USER ACTION - Choose ONE of the following:
Option 1: Install Context7 MCP Server (Recommended)
Please install the context7 MCP server in your coding assistant to access the latest Strands Evals SDK documentation. Installation steps vary by assistant:
@upstash/context7-mcpNote: If you're unsure how to install MCP servers in your coding assistant, please consult your assistant's support resources or choose Option 2 below (clone source code).
After installation, you'll be able to query: "Get documentation for strands-agents-evals focusing on Case, Experiment, and Evaluator classes"
Option 2: Clone Strands Evals SDK Source Code
If you cannot install context7 MCP or prefer to work with source code directly:
cd <your-evaluation-project>
git clone https://github.com/strands-agents/evals strands-agents-evals-sourceIMPORTANT: You MUST NOT proceed with implementation until the user has completed one of these options. Do NOT attempt to implement evaluation code using only the reference examples in Appendix C, as they may be outdated.
After the user confirms they've completed one of the above options:
If Context7 was installed:
If source code was cloned:
strands-agents-evals-source/src/strands_evals/strands-agents-evals-source/examples/Core Components:
Check Existing Requirements: Verify requirements.txt exists in repository root
# Check if requirements.txt exists
ls requirements.txtAdd Strands Evals SDK Dependencies: Update existing requirements.txt with Strands evaluation dependencies
# Add Strands Evals SDK and related dependencies
grep -q "strands-agents-evals" requirements.txt || echo "strands-agents-evals" >> requirements.txt
# Add other evaluation-specific dependencies as needed based on evaluation planInstallation: Use uv for dependency management
uv venv
source .venv/bin/activate
uv pip install -r requirements.txtConstraints:
When to Trigger: User requests results analysis or mentions generating a report
User Intent Recognition:
Execution Flow:
Parse user request: Extract any specific analysis focus (e.g., "focus on failures", "prioritize critical issues")
Navigate to evaluation project directory:
cd <your-evaluation-project> # Navigate to the directory containing both agent folder and eval/Load and analyze the evaluation results from eval/results/
Follow this execution flow:
Results Analysis Process:
a. Data Validation: Ensure results are from real execution:
b. Results Analysis: Analyze evaluation outcomes:
c. Insights Generation: Identify key findings:
Improvement Recommendations: Generate specific, actionable recommendations:
a. Prioritized Recommendations: Based on evaluation findings:
Critical Issues (Immediate attention required)
Quality Improvements (Medium-term enhancements)
Enhancement Opportunities (Future improvements)
b. Evidence-Based Recommendations: All recommendations must cite specific data:
Advisory Report Generation: Create focused report using the template structure (see Appendix B: Evaluation Report Template) with:
IMPORTANT: Follow all HTML comment instructions (<!-- ACTION REQUIRED: ... -->) in the template when generating content, then remove these comment instructions from the final report - they are template guidance only and should not appear in the generated report.
Report completion with key findings and ask: "Would you like me to help implement any of these recommendations?"
Always check for these indicators of simulated results:
Good Recommendations:
Poor Recommendations:
Ensure your advisory report:
Evaluation Report Template: See Appendix B: Evaluation Report Template
Constraints:
Finalize the evaluation and prepare deliverables.
Constraints:
<your-evaluation-project>/ # Your chosen project name
├── <your-agent-folder>/ # Your chosen agent folder name
│ └── [agent source code]
└── eval/
├── eval-plan.md
├── test-cases.jsonl
├── results/
├── run_evaluation.py
├── eval-report.md
└── README.mdagent_path: "./chatbot-agent"
evaluation_focus: "response quality and tool calling accuracy"Complete Evaluation Flow:
Phase 1 - Planning:
User: "I need to evaluate my customer support chatbot at ./chatbot-agent. Focus on response quality and tool calling accuracy."
Assistant: "I'll create an evaluation plan for your customer support chatbot..."
[Creates eval/eval-plan.md with 2 key metrics and 3 test scenarios]
Phase 2 - Data Generation:
User: "Yes, generate 5 test cases"
Assistant: "I'll generate 5 test cases covering the scenarios..."
[Creates eval/test-cases.jsonl with 2 basic queries, 2 tool-calling scenarios, 1 edge case]
Phase 3 - Evaluation Execution:
User: "Run the evaluation"
Assistant: "I'll implement and execute the evaluation using Strands Evals SDK..."
[Creates eval/run_evaluation.py, runs evaluation]
Results: Overall success rate: 80%, Response Quality: 4.2/5, Tool Call Accuracy: 75%
Phase 4 - Analysis:
User: "Yes, analyze the results"
Assistant: "I'll analyze the evaluation results and generate recommendations..."
[Creates eval/eval-report.md]
Key findings: Strong performance on basic queries (100% success), Tool calling needs improvement (25% failure rate)User: "Create an evaluation plan for my agent at ./my-agent"
Assistant: [Creates initial plan in eval/eval-plan.md]
User: "Add more focus on error handling"
Assistant: "I'll update the evaluation plan to include error handling metrics..."
[Updates eval/eval-plan.md]
User: "Generate test cases with more edge cases"
Assistant: "I'll generate test cases with additional edge case coverage..."
[Updates eval/test-cases.jsonl]After running all phases, your agent repository will have the following structure:
<your-evaluation-project>/ # Your chosen project name (e.g., my-chatbot-eval)
├── <your-agent-folder>/ # Your chosen agent folder name (e.g., chatbot-agent)
│ └── [existing agent files]
└── eval/ # All evaluation files (sibling to agent folder)
├── eval-plan.md # Complete evaluation specification and plan
├── test-cases.jsonl # Generated test scenarios
├── README.md # Running instructions and usage examples
├── run_evaluation.py # Strands Evals SDK evaluation implementation
├── results/ # Evaluation outputs
│ └── [timestamp]/ # Timestamped evaluation results
└── eval-report.md # Analysis and recommendationsNote:
<your-evaluation-project>, <your-agent-folder>) are placeholders - use any names that fit your projectEvalKit automatically manages phase dependencies:
If a user requests a phase without prerequisites:
Example: User says "run the evaluation" but no test cases exist
Response: "I don't see any test cases yet. Would you like me to:
After completing each phase, suggest the logical next step:
Issue: User requests evaluation but no agent path provided
Issue: Evaluation plan doesn't exist when user requests test generation
Issue: Test cases don't exist when user requests evaluation
Issue: Test data generation fails
python -m json.tool < eval/test-cases.jsonlIssue: Evaluation implementation fails with Strands Evals SDK errors
Issue: Import errors for evaluation dependencies
uv pip install -r requirements.txtsource .venv/bin/activateIssue: Agent execution fails during evaluation
Issue: User is unsure what to do next
The following template is used for creating eval-plan.md:
# Evaluation Plan for [AGENT NAME]
## 1. Evaluation Requirements
<!--
ACTION REQUIRED: User input and interpreted evaluation requirements. Defaults to 1-2 key metrics if unspecified.
-->
- **User Input:** `"$ARGUMENTS"` or "No Input"
- **Interpreted Evaluation Requirements:** [Parsed from user input - highest priority]
---
## 2. Agent Analysis
| **Attribute** | **Details** |
| :-------------------- | :---------------------------------------------------------- |
| **Agent Name** | [Agent name] |
| **Purpose** | [Primary purpose and use case in 1-2 sentences] |
| **Core Capabilities** | [Key functionalities the agent provides] |
| **Input** | [Short description, Data types, schemas] |
| **Output** | [Short description, Response types, schemas] |
| **Agent Framework** | [e.g., CrewAI, LangGraph, AutoGen, Custom/None] |
| **Technology Stack** | [Programming language, frameworks, libraries, dependencies] |
**Agent Architecture Diagram:**
[Mermaid diagram illustrating:
- Agent components and their relationships
- Data flow between components
- External integrations (APIs, databases, tools)
- User interaction points]
**Key Components:**
- **[Component Name 1]:** [Brief description of purpose and functionality]
- **[Component Name 2]:** [Brief description of purpose and functionality]
- [Additional components as needed]
**Available Tools:**
- **[Tool Name 1]:** [Purpose and usage]
- **[Tool Name 2]:** [Purpose and usage]
- [Additional tools as needed]
**Observability Status**
- **Tracing Framework** [Fully/Partially/Not Instrumented, Framework name, version]
- **Custom Attributes** [Yes/No, Key custom attributes if present]
---
## 3. Evaluation Metrics
<!--
ACTION REQUIRED: If no specific user requirements are provided, use a minimal number of metrics (1-2 metrics) focusing on the most critical aspects of agent performance.
-->
### [Metric Name 1]
- **Evaluation Area:** [Final response quality/tool call accuracy/...]
- **Description:** [What is measured and why]
- **Method:** [Code-based | LLM-as-Judge ]
### [Metric Name 2]
[Repeat for each metric]
---
## 4. Test Data Generation
<!--
ACTION REQUIRED: Keep scenarios minimal and focused. Do not propose more than 3 scenarios.
-->
- **[Test Scenario 1]**: [Description and purpose, complexity]
- **[Test scenario 2]**: [Description and purpose, complexity]
- **Total number of test cases**: [SHOULD NOT exceed 3]
---
## 5. Evaluation Implementation Design
### 5.1 Evaluation Code Structure
<!--
ACTION REQUIRED: The code structure below will be adjusted based on your evaluation requirements and existing agent codebase. This is the recommended starting structure. Only adjust it if necessary.
-->
./ # Repository root directory
├── requirements.txt # Consolidated dependencies
├── .venv/ # Python virtual environment (created by uv)
│
└── eval/ # Evaluation workspace
├── README.md # Running instructions and usage examples (always present)
├── run_evaluation.py # Strands Evals SDK evaluation implementation (always present)
├── results/ # Evaluation outputs (always present)
├── eval-plan.md # This evaluation specification and plan (always present)
└── test-cases.jsonl # Generated test cases (from evalkit.data)
### 5.2 Recommended Evaluation Technical Stack
| **Component** | **Selection** |
| :----------------------- | :------------------------------------------------------------ |
| **Language/Version** | [e.g., Python 3.11, Node.js 18+] |
| **Evaluation Framework** | [Strands Evals SDK (default)] |
| **Evaluators** | [OutputEvaluator, TrajectoryEvaluator, InteractionsEvaluator] |
| **Agent Integration** | [e.g., Direct import, API] |
| **Results Storage** | [e.g., JSON files (default)] |
---
## 6. Progress Tracking
### 6.1 User Requirements Log
| **Timestamp** | **Phase** | **Requirement** |
| :----------------- | :-------- | :------------------------------------------------------------------- |
| [YYYY-MM-DD HH:MM] | Planning | [User input from $ARGUMENTS, or "No specific requirements provided"] |
### 6.2 Evaluation Progress
| **Timestamp** | **Component** | **Status** | **Notes** |
| :----------------- | :--------------- | :------------------------------ | :--------------------------------------------- |
| [YYYY-MM-DD HH:MM] | [Component name] | [In Progress/Completed/Blocked] | [Technical details, blockers, or achievements] |The following template is used for creating eval-report.md:
# Agent Evaluation Report for [AGENT NAME]
## Executive Summary
<!--
ACTION REQUIRED: Provide high-level evaluation results and key findings. Focus on actionable insights for stakeholders.
-->
- **Test Scale**: [N] test cases
- **Success Rate**: [XX.X%]
- **Status**: [Excellent/Good/Poor]
- **Strengths**: [Specific capability or metric] [Performance highlight] [Reliability aspect]
- **Critical Issues**: [Blocking issue + impact] [Performance bottleneck] [Safety/compliance concern]
- **Action Priority**: [Critical fixes] [Improvements] [Enhancements]
---
## Evaluation Results
### Test Case Coverage
<!--
ACTION REQUIRED: List all test scenarios that were evaluated, providing context for the results.
-->
- **[Test Scenario 1]**: [Description and coverage]
- **[Test Scenario 2]**: [Description and coverage]
- [Additional scenarios as needed]
### Results
| **Metric** | **Score** | **Target** | **Status** |
| :-------------- | :-------- | :--------- | :---------- |
| [Metric Name 1] | [XX.X%] | [XX%] | [Pass/Fail] |
| [Metric Name 2] | [X.X/5] | [4.0+] | [Pass/Fail] |
| [Metric Name 3] | [XX.X%] | [95%+] | [Pass/Fail] |
### Results Summary
[Brief description of overall performance and findings across metrics]
---
## Agent Success Analysis
<!--
ACTION REQUIRED: Focus on what the agent does well. Provide specific evidence and contributing factors for successful performance.
-->
### Strengths
- **[Strength Name 1]**: [What the agent does exceptionally well]
- **Evidence**: [Specific metrics and examples]
- **Contributing Factors**: [Why this works well]
- **[Strength Name 2]**: [What the agent does exceptionally well]
- **Evidence**: [Specific metrics and examples]
- **Contributing Factors**: [Why this works well]
[Repeat pattern for additional strengths]
### High-Performing Scenarios
- **[Scenario Type 1]**: [Category of tasks where agent excels]
- **Key Characteristics**: [What makes these scenarios successful]
- **[Scenario Type 2]**: [Category of tasks where agent excels]
- **Key Characteristics**: [What makes these scenarios successful]
[Repeat pattern for additional scenarios]
---
## Agent Failure Analysis
<!--
ACTION REQUIRED: Analyze failures systematically. Provide root cause analysis and specific improvement recommendations with expected impact.
-->
### Issue 1 - [Priority Level]
- **Issue**: [Clear problem statement with evaluation metrics]
- **Root Cause**: [Technical analysis of why this occurred — path/to/file.py:START-END]
- **Evidence**: [Specific data points from results]
- **Impact**: [Effect on overall performance]
- **Priority Fixes**:
- P1 — [Fix name]: [One-line solution] → Expected gain: [Metric +X]
- P2 — [Fix name]: [One-line solution] → Expected gain: [Metric improvement]
### Issue 2 - [Priority Level]
[Repeat structure for additional issues]
---
## Action Items & Recommendations
<!--
ACTION REQUIRED: Provide specific, implementable tasks with clear steps. Prioritize by impact and effort required.
-->
### [Item Name] - Priority [Number] ([Critical/Enhancement])
- **Description**: [Description of this item]
- **Actions**:
- [ ] [Specific task with implementation steps]
- [ ] [Specific task with implementation steps]
- [ ] [Additional tasks as needed]
### [Additional Item Name] - Priority [Number] ([Critical/Enhancement])
[Repeat structure for additional action items]
---
## Artifacts & Reproduction
### Reference Materials
- **Agent Code**: [Path to agent implementation]
- **Test Cases**: [Path to test cases]
- **Traces**: [Path to traces]
- **Results**: [Path to results files]
- **Evaluation Code**: [Path to evaluation implementation]
---
## Evaluation Limitations and Improvement
<!--
ACTION REQUIRED: Identify limitations in the current evaluation approach and suggest improvements for future iterations.
-->
### Test Data Improvement
- **Current Limitations**: [Evaluation scope limitations]
- **Recommended Improvements**: [Specific suggestions for test data enhancement]
### Evaluation Code Enhancement
- **Current Limitations**: [Limitations of evaluation implementation and metrics]
- **Recommended Improvements**: [Specific suggestions for evaluation code improvement]
### [Additional Improvement Area]
[Repeat structure for other evaluation improvement areas]© mikeyobrien, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/eval of mikeyobrien/rho.
Open the folder on GitHubat commit 073a3ee
Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval this skillmikeyobrien/rho | 371 | — | ~9.7k | Automated safety check: Pass | MIT | |
| Eval Guidemicrosoft/eval-guide | 138 | — | ~22k | Automated safety check: Warn | MIT | |
| Skill Testdatabricks-solutions/ai-dev-kit | 1.9k | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Eval Triage And Improvementmicrosoft/eval-guide | 138 | — | ~5.9k | Automated safety check: Pass | MIT | |
| Testing Livekit Agentslivekit-examples/agent-starter-python | 264 | 1 repos | ~1.9k | Automated safety check: Pass | MIT | |
| Emcaklofas/kicad-happy | 1.3k | 1 repos | ~2.8k | Automated safety check: Pass | MIT |
microsoft/eval-guide
Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.
databricks-solutions/ai-dev-kit
Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit.
microsoft/eval-guide
A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps…
livekit-examples/agent-starter-python
Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js).
aklofas/kicad-happy
EMC pre-compliance risk analysis for KiCad PCB designs — 18 check categories, 44 rule IDs covering ground planes, decoupling, I/O filtering, switching harmonics, clock routing, differential pair…
swig/swig
Run SWIG test suite for specific languages. An agent skill from swig/swig.
mikeyobrien/rho
Register an agent email address on Rhobot Mail (name@rhobot.dev).
mikeyobrien/rho
Install and configure Rho from scratch (Doom-style init.toml + sync).
mikeyobrien/rho
Open URLs and launch apps on Android. An agent skill from mikeyobrien/rho.
mikeyobrien/rho
Keep CHANGELOG.md idiomatic (Keep a Changelog) and cut a tag-based GitHub release that triggers npm publish CI.
mikeyobrien/rho
Manage agent email at name@rhobot.dev via the Rhobot Mail API.
mikeyobrien/rho
Create Tasker profiles and tasks via XML for Android automation.
Categories
Plan and run conversational AI agent evaluations with test generation and analysis. Eval is an agent skill from mikeyobrien/rho. Plan and run conversational AI agent evaluations with test generation and analysis.
Eval fits situations like: tasks that involve Test generation; tasks that involve Agent evaluation and testing; tasks that involve Chatbots and conversational support.
Run `npx skills add mikeyobrien/rho --skill eval -a claude-code`. Or copy the skill folder (skills/eval in mikeyobrien/rho) into .claude/skills/eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mikeyobrien/rho --skill eval -a codex`. Or copy the skill folder (skills/eval in mikeyobrien/rho) into .agents/skills/eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mikeyobrien/rho --skill eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval, .gemini/skills/eval, .github/skills/eval and .opencode/skills/eval in your project.
Going by SKILL.md and its folder, Eval needs the command-line tools its instructions call (uv, git and python). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 9.7k tokens (SKILL.md is roughly 39k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval: Eval Guide (microsoft/eval-guide, 138 stars), Skill Test (databricks-solutions/ai-dev-kit, 1.9k stars), Eval Triage And Improvement (microsoft/eval-guide, 138 stars) and Testing Livekit Agents (livekit-examples/agent-starter-python, 264 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mikeyobrien (a GitHub user) maintains it in mikeyobrien/rho, which has 371 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 1, 2026.
Source: mikeyobrien/rho on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.