Arize Evaluator
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
Create evaluator functions in evaluators.ts for Output SDK workflows.
$ npx skills add growthxai/output --skill output-dev-evaluator-function -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install growthxai/output output-dev-evaluator-function --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .claude/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function .claude/skills/output-dev-evaluator-function && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "output-dev-evaluator-function" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function into .claude/skills/output-dev-evaluator-function/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-evaluator-function", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-functionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add growthxai/output --skill output-dev-evaluator-function -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install growthxai/output output-dev-evaluator-function --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .agents/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function .agents/skills/output-dev-evaluator-function && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "output-dev-evaluator-function" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function into .agents/skills/output-dev-evaluator-function/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-evaluator-function", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add growthxai/output --skill output-dev-evaluator-function -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install growthxai/output output-dev-evaluator-function --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function .cursor/skills/output-dev-evaluator-function && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "output-dev-evaluator-function" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function into .cursor/skills/output-dev-evaluator-function/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-evaluator-function", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/growthxai/output.git --path coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add growthxai/output --skill output-dev-evaluator-function -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install growthxai/output output-dev-evaluator-function --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function .gemini/skills/output-dev-evaluator-function && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "output-dev-evaluator-function" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function into .gemini/skills/output-dev-evaluator-function/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-evaluator-function", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install growthxai/output output-dev-evaluator-functionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add growthxai/output --skill output-dev-evaluator-function -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .github/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function .github/skills/output-dev-evaluator-function && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "output-dev-evaluator-function" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function into .github/skills/output-dev-evaluator-function/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-evaluator-function", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add growthxai/output --skill output-dev-evaluator-function -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install growthxai/output output-dev-evaluator-function --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function .opencode/skills/output-dev-evaluator-function && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "output-dev-evaluator-function" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function into .opencode/skills/output-dev-evaluator-function/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-evaluator-function", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
output-dev-evaluator-functionCreate evaluator functions in evaluators.ts for Output SDK workflows.
Output Dev Evaluator Function is an agent skill from growthxai/output. Create evaluator functions in evaluators.ts for Output SDK workflows. Use when implementing quality assessment, validation logic, or content evaluation.
Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already… The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 52b51ac. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Output Dev Evaluator Function loads about 5.5k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 643 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from growthxai/output at commit 52b51ac, republished under its Apache-2.0 licence (© growthxai). 643 words, ~5,458 tokens.
.claude/skills/output-dev-evaluator-function/SKILL.md (or your agent's skills folder).This skill documents how to create evaluator functions in evaluators.ts for Output SDK workflows. Evaluators are used to assess quality, validate outputs, and provide confidence-scored judgments about workflow results.
For smaller workflows, use a single evaluators.ts file:
src/workflows/{workflow-name}/
├── workflow.ts
├── steps.ts
├── evaluators.ts # All evaluators in one file
├── types.ts
└── ...For larger workflows with many evaluators, use an evaluators/ folder:
src/workflows/{workflow-name}/
├── workflow.ts
├── steps.ts
├── evaluators/ # Evaluators split into individual files
│ ├── quality.ts
│ ├── accuracy.ts
│ └── completeness.ts
├── types.ts
└── ...Important: evaluator() calls MUST be in files containing 'evaluators' in the path:
src/workflows/my_workflow/evaluators.ts ✓src/workflows/my_workflow/evaluators/quality.ts ✓src/shared/evaluators/common_evaluators.ts ✓src/workflows/my_workflow/helpers.ts ✗ (cannot contain evaluator() calls)Evaluators are Temporal activities with strict import rules to ensure deterministic replay.
./utils.js, ./types.js, ./helpers.js./lib/helpers.js../../shared/utils/*.js../../shared/clients/*.js../../shared/services/*.jsExample of WRONG imports:
// WRONG - evaluators cannot import other evaluators
import { otherEvaluator } from '../../shared/evaluators/other.js'; // ✗
import { anotherEvaluator } from './other_evaluators.js'; // ✗// CORRECT - Import from @outputai/core
import {
evaluator,
z,
EvaluationBooleanResult,
EvaluationNumberResult,
EvaluationStringResult,
EvaluationFeedback
} from '@outputai/core';
// WRONG - Never import z from zod
import { z } from 'zod';// CORRECT - Use @outputai/llm wrapper
import { generateText, aiSdk } from '@outputai/llm';
// WRONG - Never call LLM providers directly
import OpenAI from 'openai';All imports MUST use .js extension:
// CORRECT
import { BlogContent } from './types.js';
// WRONG - Missing .js extension
import { BlogContent } from './types';import { evaluator, z, EvaluationBooleanResult } from '@outputai/core';
export const myEvaluator = evaluator( {
name: 'my_evaluator',
description: 'Description of what this evaluator assesses',
inputSchema: z.object( { /* input schema */ } ),
fn: async input => {
// Evaluation logic
return new EvaluationBooleanResult( {
value: true,
confidence: 0.95
} );
}
} );Unique identifier for the evaluator. Use snake_case.
name: 'evaluate_content_quality'Human-readable description of what the evaluator assesses.
description: 'Evaluate the quality and completeness of generated content'Schema for validating evaluator input.
inputSchema: z.object( {
content: z.string(),
expectedLength: z.number()
} )The evaluator execution function. Returns an evaluation result with value and confidence.
fn: async input => {
const isValid = input.content.length >= input.expectedLength;
return new EvaluationBooleanResult( {
value: isValid,
confidence: 0.95
} );
}Use for pass/fail or true/false evaluations:
import { EvaluationBooleanResult } from '@outputai/core';
return new EvaluationBooleanResult( {
value: true, // boolean result
confidence: 0.95, // 0.0 to 1.0
reasoning: 'Optional explanation of the evaluation'
} );Use for numeric scores or ratings:
import { EvaluationNumberResult } from '@outputai/core';
return new EvaluationNumberResult( {
value: 85, // numeric result (e.g., 0-100 score)
confidence: 0.85, // 0.0 to 1.0
reasoning: 'Optional explanation of the score'
} );Use for categorical or text-based evaluations:
import { EvaluationStringResult } from '@outputai/core';
return new EvaluationStringResult( {
value: 'positive', // string result (e.g., category, sentiment, label)
confidence: 0.9, // 0.0 to 1.0
reasoning: 'Optional explanation of the classification'
} );| Property | Type | Required | Description |
|---|---|---|---|
value | boolean, number, or string | Yes | The evaluation result |
confidence | number (0.0-1.0) | Yes | Confidence in the evaluation |
reasoning | string | No | Explanation of the evaluation |
name | string | No | Name for this specific result (useful in dimensions) |
feedback | EvaluationFeedback[] | No | Array of feedback objects with issues and suggestions |
dimensions | EvaluationResult[] | No | Nested results for multi-dimensional evaluation |
import { evaluator, z, EvaluationBooleanResult } from '@outputai/core';
export const evaluateCompleteness = evaluator( {
name: 'evaluate_completeness',
description: 'Check if content meets minimum length requirements',
inputSchema: z.object( {
content: z.string(),
minLength: z.number().default( 100 )
} ),
fn: async ( { content, minLength } ) => {
const isComplete = content.length >= minLength;
return new EvaluationBooleanResult( {
value: isComplete,
confidence: 1.0,
reasoning: isComplete ?
`Content has ${content.length} characters, meets minimum of ${minLength}` :
`Content has ${content.length} characters, below minimum of ${minLength}`
} );
}
} );import { evaluator, z, EvaluationBooleanResult } from '@outputai/core';
export const evaluateGibberish = evaluator( {
name: 'evaluate_gibberish',
description: 'Check if a given string is gibberish',
inputSchema: z.string(),
fn: async content => {
const gibberishPatterns = [ 'foo', 'bar', 'lorem', 'ipsum' ];
const isGibberish = gibberishPatterns.some( p => content.toLowerCase().includes( p ) );
return new EvaluationBooleanResult( {
value: !isGibberish,
confidence: 0.95
} );
}
} );import { evaluator, z, EvaluationNumberResult } from '@outputai/core';
export const evaluateReadability = evaluator( {
name: 'evaluate_readability',
description: 'Calculate readability score based on sentence structure',
inputSchema: z.object( {
content: z.string()
} ),
fn: async ( { content } ) => {
const sentences = content.split( /[.!?]+/ ).filter( s => s.trim() );
const words = content.split( /\s+/ ).filter( w => w.trim() );
const avgWordsPerSentence = words.length / Math.max( sentences.length, 1 );
// Simple readability score (lower avg words = more readable)
const score = Math.max( 0, Math.min( 100, 100 - ( avgWordsPerSentence - 15 ) * 5 ) );
return new EvaluationNumberResult( {
value: Math.round( score ),
confidence: 0.8,
reasoning: `Average ${avgWordsPerSentence.toFixed( 1 )} words per sentence`
} );
}
} );import { evaluator, z, EvaluationStringResult } from '@outputai/core';
export const evaluateSentiment = evaluator( {
name: 'evaluate_sentiment',
description: 'Classify the sentiment of content',
inputSchema: z.object( {
content: z.string()
} ),
fn: async ( { content } ) => {
const positiveWords = [ 'great', 'excellent', 'amazing', 'good', 'love' ];
const negativeWords = [ 'bad', 'terrible', 'awful', 'hate', 'poor' ];
const lowerContent = content.toLowerCase();
const positiveCount = positiveWords.filter( w => lowerContent.includes( w ) ).length;
const negativeCount = negativeWords.filter( w => lowerContent.includes( w ) ).length;
const { sentiment, confidence } = positiveCount > negativeCount ?
{ sentiment: 'positive', confidence: Math.min( 0.95, 0.6 + positiveCount * 0.1 ) } :
negativeCount > positiveCount ?
{ sentiment: 'negative', confidence: Math.min( 0.95, 0.6 + negativeCount * 0.1 ) } :
{ sentiment: 'neutral', confidence: 0.7 };
return new EvaluationStringResult( {
value: sentiment,
confidence,
reasoning: `Found ${positiveCount} positive and ${negativeCount} negative indicators`
} );
}
} );Note: Evaluators are self-contained components that don't share schemas across steps, so defining aiSdk.Output.object() schemas inline is acceptable here. For workflow steps that share schemas, define them in types.ts instead.
generateText arguments: prompt, promptDir, variables, tools, output, toolChoice, stopWhen, abortSignal.
import { evaluator, z, EvaluationNumberResult } from '@outputai/core';
import { generateText, aiSdk } from '@outputai/llm';
export const evaluateSignalToNoise = evaluator( {
name: 'evaluate_signal_to_noise',
description: 'Evaluate the signal-to-noise ratio of content',
inputSchema: z.object( {
title: z.string(),
content: z.string()
} ),
fn: async ( { title, content } ) => {
const { output } = await generateText( {
prompt: 'signal_noise@v1', // References prompts/signal_noise@v1.prompt
variables: {
title,
content
},
output: aiSdk.Output.object( {
schema: z.object( {
score: z.number().describe( 'Signal-to-noise score 0-100' )
} )
} )
} );
return new EvaluationNumberResult( {
value: output.score,
confidence: 0.85
} );
}
} );import { evaluator, z, EvaluationBooleanResult } from '@outputai/core';
import { generateText, aiSdk } from '@outputai/llm';
export const evaluateFactualAccuracy = evaluator( {
name: 'evaluate_factual_accuracy',
description: 'Check if content contains factual claims that can be verified',
inputSchema: z.object( {
content: z.string(),
topic: z.string()
} ),
fn: async ( { content, topic } ) => {
const { output } = await generateText( {
prompt: 'factual_check@v1',
variables: { content, topic },
output: aiSdk.Output.object( {
schema: z.object( {
isFactual: z.boolean().describe( 'Whether content appears factually accurate' ),
confidence: z.number().describe( 'Confidence in assessment 0-1' ),
issues: z.array( z.string() ).optional().describe( 'Any factual issues found' )
} )
} )
} );
return new EvaluationBooleanResult( {
value: output.isFactual,
confidence: output.confidence,
reasoning: output.issues?.length ?
`Issues found: ${output.issues.join( ', ' )}` :
'No factual issues detected'
} );
}
} );import { evaluator, z, EvaluationStringResult } from '@outputai/core';
import { generateText, aiSdk } from '@outputai/llm';
export const evaluateContentCategory = evaluator( {
name: 'evaluate_content_category',
description: 'Classify content into a category',
inputSchema: z.object( {
content: z.string(),
categories: z.array( z.string() )
} ),
fn: async ( { content, categories } ) => {
const { output } = await generateText( {
prompt: 'categorize_content@v1',
variables: {
content,
categories: categories.join( ', ' )
},
output: aiSdk.Output.object( {
schema: z.object( {
category: z.string().describe( 'The best matching category' ),
confidence: z.number().describe( 'Confidence in classification 0-1' ),
explanation: z.string().describe( 'Why this category was chosen' )
} )
} )
} );
return new EvaluationStringResult( {
value: output.category,
confidence: output.confidence,
reasoning: output.explanation
} );
}
} );Use the feedback field to provide actionable improvement suggestions alongside your evaluation result. Import EvaluationFeedback from @outputai/core to create feedback objects.
import { evaluator, z, EvaluationStringResult, EvaluationFeedback } from '@outputai/core';
export const evaluateWithFeedback = evaluator( {
name: 'evaluate_with_feedback',
description: 'Evaluate content quality and provide actionable feedback',
inputSchema: z.string(),
fn: async response => {
const feedback = [];
if ( response.length < 50 ) {
feedback.push( new EvaluationFeedback( {
issue: 'Response is too short',
suggestion: 'Expand the response with more detail',
priority: 'medium'
} ) );
}
return new EvaluationStringResult( {
value: feedback.length === 0 ? 'good' : 'needs_improvement',
confidence: 0.85,
feedback: feedback
} );
}
} );| Property | Type | Description |
|---|---|---|
issue | string | The problem identified |
suggestion | string | Recommended fix |
priority | string | Priority level (e.g., 'low', 'medium', 'high') |
Use the dimensions field to nest EvaluationResult instances for sub-scores. Each dimension should use the name field to identify it.
import { evaluator, z, EvaluationStringResult, EvaluationNumberResult } from '@outputai/core';
export const evaluateMultiDimensional = evaluator( {
name: 'evaluate_multi_dimensional',
description: 'Evaluate content across multiple quality dimensions',
inputSchema: z.string(),
fn: async response => {
const coherenceScore = calculateCoherence( response );
const relevanceScore = calculateRelevance( response );
const overallScore = ( coherenceScore + relevanceScore ) / 2;
return new EvaluationStringResult( {
value: overallScore > 0.7 ? 'high_quality' : 'low_quality',
confidence: 0.9,
dimensions: [
new EvaluationNumberResult( {
value: coherenceScore,
confidence: 0.85,
name: 'coherence'
} ),
new EvaluationNumberResult( {
value: relevanceScore,
confidence: 0.88,
name: 'relevance'
} )
]
} );
}
} );Based on a real workflow evaluator file:
import { evaluator, z, EvaluationBooleanResult, EvaluationNumberResult } from '@outputai/core';
import { generateText, aiSdk } from '@outputai/llm';
import { blogContentSchema } from './types.js';
import type { BlogContent, QualityMetrics } from './types.js';
// Simple boolean evaluator
export const evaluateMinimumLength = evaluator( {
name: 'evaluate_minimum_length',
description: 'Check if blog content meets minimum length requirements',
inputSchema: blogContentSchema,
fn: async ( input: BlogContent ) => {
const MIN_TOKENS = 500;
const meetsRequirement = input.tokenCount >= MIN_TOKENS;
return new EvaluationBooleanResult( {
value: meetsRequirement,
confidence: 1.0,
reasoning: `Content has ${input.tokenCount} tokens (minimum: ${MIN_TOKENS})`
} );
}
} );
// LLM-powered number evaluator
export const evaluateSignalToNoise = evaluator( {
name: 'evaluate_signal_to_noise',
description: 'Evaluate the signal-to-noise ratio of blog content',
inputSchema: blogContentSchema,
fn: async ( input: BlogContent ) => {
const { output } = await generateText( {
prompt: 'signal_noise@v1',
variables: {
title: input.title,
content: input.content
},
output: aiSdk.Output.object( {
schema: z.object( {
score: z.number().describe( 'Signal-to-noise score 0-100' )
} )
} )
} );
return new EvaluationNumberResult( {
value: output.score,
confidence: 0.85
} );
}
} );
// LLM-powered boolean evaluator
export const evaluateRelevance = evaluator( {
name: 'evaluate_relevance',
description: 'Check if content is relevant to the stated topic',
inputSchema: z.object( {
content: z.string(),
topic: z.string(),
keywords: z.array( z.string() )
} ),
fn: async ( { content, topic, keywords } ) => {
const { output } = await generateText( {
prompt: 'relevance_check@v1',
variables: { content, topic, keywords: keywords.join( ', ' ) },
output: aiSdk.Output.object( {
schema: z.object( {
isRelevant: z.boolean(),
relevanceScore: z.number().describe( 'Relevance score 0-1' ),
explanation: z.string()
} )
} )
} );
return new EvaluationBooleanResult( {
value: output.isRelevant,
confidence: output.relevanceScore,
reasoning: output.explanation
} );
}
} );// Boolean for pass/fail decisions
return new EvaluationBooleanResult( { value: true, confidence: 0.9 } );
// Number for scores and ratings
return new EvaluationNumberResult( { value: 85, confidence: 0.85 } );
// String for categories, labels, or classifications
return new EvaluationStringResult( { value: 'positive', confidence: 0.9 } );// High confidence for deterministic checks
confidence: 1.0 // e.g., length checks, pattern matching
// Medium confidence for heuristic-based evaluations
confidence: 0.85 // e.g., LLM-based assessments
// Lower confidence for uncertain evaluations
confidence: 0.7 // e.g., subjective quality judgmentsreturn new EvaluationBooleanResult( {
value: false,
confidence: 0.95,
reasoning: `Content contains ${errorCount} grammatical errors, exceeding threshold of ${maxErrors}`
} );// Good - single responsibility
export const evaluateGrammar = evaluator( { ... } );
export const evaluateReadability = evaluator( { ... } );
export const evaluateTone = evaluator( { ... } );
// Avoid - doing too much in one evaluator
export const evaluateEverything = evaluator( { ... } );// Good - clear what is being evaluated
name: 'evaluate_content_originality'
name: 'evaluate_factual_accuracy'
name: 'evaluate_sentiment_alignment'
// Avoid - vague names
name: 'check'
name: 'validate'
name: 'evaluate_stuff'feedback: [
new EvaluationFeedback( {
issue: 'Missing conclusion paragraph',
suggestion: 'Add a summary paragraph at the end',
priority: 'high'
} )
]dimensions: [
new EvaluationNumberResult( { value: 8, confidence: 0.9, name: 'coherence' } ),
new EvaluationNumberResult( { value: 6, confidence: 0.85, name: 'relevance' } )
]evaluator, z, result types imported from @outputai/coregenerateText and aiSdk imported from @outputai/llm if using LLM (not direct provider).describe() instead of .min()/.max() on z.number().js extensionname, description, inputSchema, fnsnake_caseEvaluationBooleanResult, EvaluationNumberResult, or EvaluationStringResult)EvaluationFeedback imported from @outputai/core when using feedbackissue, suggestion, and priorityname field to identify sub-evaluationsoutput-dev-workflow-function - Orchestrating evaluators in workflow.tsoutput-dev-step-function - Creating step functionsoutput-dev-types-file - Defining evaluator input schemasoutput-dev-prompt-file - Creating prompt files for LLM-powered evaluatorsoutput-dev-folder-structure - Understanding project layoutoutput-eval-error-analysis — Identify what to evaluate before writing evaluators© growthxai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function of growthxai/output.
Open the folder on GitHubat commit 52b51ac
Output Dev Evaluator Function next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Output Dev Evaluator Function this skillgrowthxai/output | 440 | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | |
| Arize Evaluatorgithub/awesome-copilot | 40k | 2 repos | ~8.1k | Automated safety check: Notes | MIT | |
| LLM Evaluationdavila7/claude-code-templates | 32k | 13 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Agent Evaluationsickn33/agentic-awesome-skills | 47k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| EvaluatorsArize-ai/phoenix | 12k | — | ~1.7k | Automated safety check: Pass | Custom licence | |
| Agent Evaluation Reportingsickn33/agentic-awesome-skills | 47k | 1 repos | ~2.1k | Automated safety check: Pass | MIT |
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
davila7/claude-code-templates
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
sickn33/agentic-awesome-skills
Evaluate agent behavior with versioned cases and explicit verifiers.
Arize-ai/phoenix
Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.
sickn33/agentic-awesome-skills
A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.
PostHog/posthog
Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified.
growthxai/output
Zod schema constraints that Anthropic rejects or silently ignores when sent as structured-output tool definitions via aiSdk.Output.object().
growthxai/output
Implement an Output SDK workflow from a plan document. An agent skill from growthxai/output.
growthxai/output
View, edit, and set encrypted credentials in an Output.ai project.
growthxai/output
Wire encrypted credentials to environment variables using the credential: convention.
growthxai/output
Initialize encrypted credentials for an Output.ai project. An agent skill from growthxai/output.
growthxai/output
Debug Output SDK workflow issues. An agent skill from growthxai/output.
Create evaluator functions in evaluators.ts for Output SDK workflows. Output Dev Evaluator Function is an agent skill from growthxai/output.ts for Output SDK workflows.
Output Dev Evaluator Function fits situations like: implementing quality assessment; validation logic; content evaluation.
Run `npx skills add growthxai/output --skill output-dev-evaluator-function -a claude-code`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function in growthxai/output) into .claude/skills/output-dev-evaluator-function in your project. Claude Code loads it when a task matches its description.
Run `npx skills add growthxai/output --skill output-dev-evaluator-function -a codex`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function in growthxai/output) into .agents/skills/output-dev-evaluator-function in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add growthxai/output --skill output-dev-evaluator-function -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/output-dev-evaluator-function, .gemini/skills/output-dev-evaluator-function, .github/skills/output-dev-evaluator-function and .opencode/skills/output-dev-evaluator-function in your project.
SKILL.md names no scripts, command-line tools or credentials: Output Dev Evaluator Function is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Write, Edit.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Output Dev Evaluator Function is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Output Dev Evaluator Function: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars) and Evaluators (Arize-ai/phoenix, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
growthxai (a GitHub organization) maintains it in growthxai/output, which has 440 GitHub stars. The repository holds 52 skills in this directory. The repository was last updated on October 7, 2026.
Source: growthxai/output on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.