Agent skill

Output Dev Evaluator Function

by growthxai in growthxai/output

Create evaluator functions in evaluators.ts for Output SDK workflows.

Apache-2.0Auto-check passed

Install Output Dev Evaluator Function

skills CLI
$ npx skills add growthxai/output --skill output-dev-evaluator-function -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install growthxai/output output-dev-evaluator-function --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .claude/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function .claude/skills/output-dev-evaluator-function && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
output-dev-evaluator-function
GitHub stars
440
Token cost
~5.5k tokens
SKILL.md length
643 words
Files
1
Skills in repo
52
Repo updated
First seen
Licence
Apache-2.0

At a glance

Create evaluator functions in evaluators.ts for Output SDK workflows.

  • Works in 7 steps: Use Appropriate Result Types → Provide Meaningful Confidence Scores → Include Reasoning for Transparency → …
  • Implementing quality assessment
  • SKILL.md covers Overview, When to Use This Skill, File Organization and Component Location Rules, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Output Dev Evaluator Function is an agent skill from growthxai/output. Create evaluator functions in evaluators.ts for Output SDK workflows. Use when implementing quality assessment, validation logic, or content evaluation.

Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already… The licence is Apache-2.0.

When your agent uses it

  • Implementing quality assessment
  • Validation logic
  • Content evaluation

Example prompts

  • “/output-dev-evaluator-function”

Requirements

  • Pre-approved tools (allowed-tools): Read, Write, Edit

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Use Appropriate Result Types
  2. Provide Meaningful Confidence Scores
  3. Include Reasoning for Transparency
  4. Keep Evaluators Focused
  5. Use Descriptive Names
  6. Use Feedback for Actionable Improvements
  7. Use Dimensions for Multi-Criteria Evaluation

What it can do on your machine

Read from SKILL.md and the folder at commit 52b51ac. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Output Dev Evaluator Function loads about 5.5k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 643 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from growthxai/output at commit 52b51ac, republished under its Apache-2.0 licence (© growthxai). 643 words, ~5,458 tokens.

Download SKILL.mdSave it as .claude/skills/output-dev-evaluator-function/SKILL.md (or your agent's skills folder).
name
output-dev-evaluator-function
description
Create evaluator functions in evaluators.ts for Output SDK workflows. Use when implementing quality assessment, validation logic, or content evaluation.
allowed-tools
Read, Write, Edit

Creating Evaluator Functions

Overview

This skill documents how to create evaluator functions in evaluators.ts for Output SDK workflows. Evaluators are used to assess quality, validate outputs, and provide confidence-scored judgments about workflow results.

When to Use This Skill

  • Implementing quality assessment for workflow outputs
  • Adding validation logic with confidence scores
  • Creating LLM-powered content evaluation
  • Building reusable evaluation components

File Organization

Option 1: Flat File (Default)

For smaller workflows, use a single evaluators.ts file:

src/workflows/{workflow-name}/
├── workflow.ts
├── steps.ts
├── evaluators.ts    # All evaluators in one file
├── types.ts
└── ...
Option 2: Folder-Based (Large workflows)

For larger workflows with many evaluators, use an evaluators/ folder:

src/workflows/{workflow-name}/
├── workflow.ts
├── steps.ts
├── evaluators/      # Evaluators split into individual files
│   ├── quality.ts
│   ├── accuracy.ts
│   └── completeness.ts
├── types.ts
└── ...

Component Location Rules

Important: evaluator() calls MUST be in files containing 'evaluators' in the path:

  • src/workflows/my_workflow/evaluators.ts ✓
  • src/workflows/my_workflow/evaluators/quality.ts ✓
  • src/shared/evaluators/common_evaluators.ts ✓
  • src/workflows/my_workflow/helpers.ts ✗ (cannot contain evaluator() calls)

Activity Isolation Constraints

Evaluators are Temporal activities with strict import rules to ensure deterministic replay.

Evaluators CAN import from:
  • Local workflow files: ./utils.js, ./types.js, ./helpers.js
  • Local subdirectories: ./lib/helpers.js
  • Shared utilities: ../../shared/utils/*.js
  • Shared clients: ../../shared/clients/*.js
  • Shared services: ../../shared/services/*.js
Evaluators CANNOT import:
  • Other evaluator files (activity isolation)
  • Step files
  • Workflow files

Example of WRONG imports:

typescript
// WRONG - evaluators cannot import other evaluators
import { otherEvaluator } from '../../shared/evaluators/other.js'; // ✗
import { anotherEvaluator } from './other_evaluators.js'; // ✗

Critical Import Patterns

Core Imports
typescript
// CORRECT - Import from @outputai/core
import {
  evaluator,
  z,
  EvaluationBooleanResult,
  EvaluationNumberResult,
  EvaluationStringResult,
  EvaluationFeedback
} from '@outputai/core';

// WRONG - Never import z from zod
import { z } from 'zod';
LLM Client Import (for LLM-powered evaluators)
typescript
// CORRECT - Use @outputai/llm wrapper
import { generateText, aiSdk } from '@outputai/llm';

// WRONG - Never call LLM providers directly
import OpenAI from 'openai';
ES Module Imports

All imports MUST use .js extension:

typescript
// CORRECT
import { BlogContent } from './types.js';

// WRONG - Missing .js extension
import { BlogContent } from './types';

Basic Structure

typescript
import { evaluator, z, EvaluationBooleanResult } from '@outputai/core';

export const myEvaluator = evaluator( {
  name: 'my_evaluator',
  description: 'Description of what this evaluator assesses',
  inputSchema: z.object( { /* input schema */ } ),
  fn: async input => {
    // Evaluation logic
    return new EvaluationBooleanResult( {
      value: true,
      confidence: 0.95
    } );
  }
} );

Required Properties

name (string)

Unique identifier for the evaluator. Use snake_case.

typescript
name: 'evaluate_content_quality'
description (string)

Human-readable description of what the evaluator assesses.

typescript
description: 'Evaluate the quality and completeness of generated content'
inputSchema (Zod schema)

Schema for validating evaluator input.

typescript
inputSchema: z.object( {
  content: z.string(),
  expectedLength: z.number()
} )
fn (async function)

The evaluator execution function. Returns an evaluation result with value and confidence.

typescript
fn: async input => {
  const isValid = input.content.length >= input.expectedLength;
  return new EvaluationBooleanResult( {
    value: isValid,
    confidence: 0.95
  } );
}

Result Types

EvaluationBooleanResult

Use for pass/fail or true/false evaluations:

typescript
import { EvaluationBooleanResult } from '@outputai/core';

return new EvaluationBooleanResult( {
  value: true,           // boolean result
  confidence: 0.95,      // 0.0 to 1.0
  reasoning: 'Optional explanation of the evaluation'
} );
EvaluationNumberResult

Use for numeric scores or ratings:

typescript
import { EvaluationNumberResult } from '@outputai/core';

return new EvaluationNumberResult( {
  value: 85,             // numeric result (e.g., 0-100 score)
  confidence: 0.85,      // 0.0 to 1.0
  reasoning: 'Optional explanation of the score'
} );
EvaluationStringResult

Use for categorical or text-based evaluations:

typescript
import { EvaluationStringResult } from '@outputai/core';

return new EvaluationStringResult( {
  value: 'positive',     // string result (e.g., category, sentiment, label)
  confidence: 0.9,       // 0.0 to 1.0
  reasoning: 'Optional explanation of the classification'
} );

Result Properties

PropertyTypeRequiredDescription
valueboolean, number, or stringYesThe evaluation result
confidencenumber (0.0-1.0)YesConfidence in the evaluation
reasoningstringNoExplanation of the evaluation
namestringNoName for this specific result (useful in dimensions)
feedbackEvaluationFeedback[]NoArray of feedback objects with issues and suggestions
dimensionsEvaluationResult[]NoNested results for multi-dimensional evaluation

Simple Evaluator Examples

Boolean Evaluator - Content Validation
typescript
import { evaluator, z, EvaluationBooleanResult } from '@outputai/core';

export const evaluateCompleteness = evaluator( {
  name: 'evaluate_completeness',
  description: 'Check if content meets minimum length requirements',
  inputSchema: z.object( {
    content: z.string(),
    minLength: z.number().default( 100 )
  } ),
  fn: async ( { content, minLength } ) => {
    const isComplete = content.length >= minLength;

    return new EvaluationBooleanResult( {
      value: isComplete,
      confidence: 1.0,
      reasoning: isComplete ?
        `Content has ${content.length} characters, meets minimum of ${minLength}` :
        `Content has ${content.length} characters, below minimum of ${minLength}`
    } );
  }
} );
Boolean Evaluator - Pattern Detection
typescript
import { evaluator, z, EvaluationBooleanResult } from '@outputai/core';

export const evaluateGibberish = evaluator( {
  name: 'evaluate_gibberish',
  description: 'Check if a given string is gibberish',
  inputSchema: z.string(),
  fn: async content => {
    const gibberishPatterns = [ 'foo', 'bar', 'lorem', 'ipsum' ];
    const isGibberish = gibberishPatterns.some( p => content.toLowerCase().includes( p ) );

    return new EvaluationBooleanResult( {
      value: !isGibberish,
      confidence: 0.95
    } );
  }
} );
Number Evaluator - Quality Score
typescript
import { evaluator, z, EvaluationNumberResult } from '@outputai/core';

export const evaluateReadability = evaluator( {
  name: 'evaluate_readability',
  description: 'Calculate readability score based on sentence structure',
  inputSchema: z.object( {
    content: z.string()
  } ),
  fn: async ( { content } ) => {
    const sentences = content.split( /[.!?]+/ ).filter( s => s.trim() );
    const words = content.split( /\s+/ ).filter( w => w.trim() );
    const avgWordsPerSentence = words.length / Math.max( sentences.length, 1 );

    // Simple readability score (lower avg words = more readable)
    const score = Math.max( 0, Math.min( 100, 100 - ( avgWordsPerSentence - 15 ) * 5 ) );

    return new EvaluationNumberResult( {
      value: Math.round( score ),
      confidence: 0.8,
      reasoning: `Average ${avgWordsPerSentence.toFixed( 1 )} words per sentence`
    } );
  }
} );
String Evaluator - Sentiment Classification
typescript
import { evaluator, z, EvaluationStringResult } from '@outputai/core';

export const evaluateSentiment = evaluator( {
  name: 'evaluate_sentiment',
  description: 'Classify the sentiment of content',
  inputSchema: z.object( {
    content: z.string()
  } ),
  fn: async ( { content } ) => {
    const positiveWords = [ 'great', 'excellent', 'amazing', 'good', 'love' ];
    const negativeWords = [ 'bad', 'terrible', 'awful', 'hate', 'poor' ];

    const lowerContent = content.toLowerCase();
    const positiveCount = positiveWords.filter( w => lowerContent.includes( w ) ).length;
    const negativeCount = negativeWords.filter( w => lowerContent.includes( w ) ).length;

    const { sentiment, confidence } = positiveCount > negativeCount ?
      { sentiment: 'positive', confidence: Math.min( 0.95, 0.6 + positiveCount * 0.1 ) } :
      negativeCount > positiveCount ?
        { sentiment: 'negative', confidence: Math.min( 0.95, 0.6 + negativeCount * 0.1 ) } :
        { sentiment: 'neutral', confidence: 0.7 };

    return new EvaluationStringResult( {
      value: sentiment,
      confidence,
      reasoning: `Found ${positiveCount} positive and ${negativeCount} negative indicators`
    } );
  }
} );

LLM-Powered Evaluator Examples

Note: Evaluators are self-contained components that don't share schemas across steps, so defining aiSdk.Output.object() schemas inline is acceptable here. For workflow steps that share schemas, define them in types.ts instead.

generateText arguments: prompt, promptDir, variables, tools, output, toolChoice, stopWhen, abortSignal.

Using generateText with aiSdk.Output.object() for Evaluation
typescript
import { evaluator, z, EvaluationNumberResult } from '@outputai/core';
import { generateText, aiSdk } from '@outputai/llm';

export const evaluateSignalToNoise = evaluator( {
  name: 'evaluate_signal_to_noise',
  description: 'Evaluate the signal-to-noise ratio of content',
  inputSchema: z.object( {
    title: z.string(),
    content: z.string()
  } ),
  fn: async ( { title, content } ) => {
    const { output } = await generateText( {
      prompt: 'signal_noise@v1',  // References prompts/signal_noise@v1.prompt
      variables: {
        title,
        content
      },
      output: aiSdk.Output.object( {
        schema: z.object( {
          score: z.number().describe( 'Signal-to-noise score 0-100' )
        } )
      } )
    } );

    return new EvaluationNumberResult( {
      value: output.score,
      confidence: 0.85
    } );
  }
} );
LLM Boolean Evaluation
typescript
import { evaluator, z, EvaluationBooleanResult } from '@outputai/core';
import { generateText, aiSdk } from '@outputai/llm';

export const evaluateFactualAccuracy = evaluator( {
  name: 'evaluate_factual_accuracy',
  description: 'Check if content contains factual claims that can be verified',
  inputSchema: z.object( {
    content: z.string(),
    topic: z.string()
  } ),
  fn: async ( { content, topic } ) => {
    const { output } = await generateText( {
      prompt: 'factual_check@v1',
      variables: { content, topic },
      output: aiSdk.Output.object( {
        schema: z.object( {
          isFactual: z.boolean().describe( 'Whether content appears factually accurate' ),
          confidence: z.number().describe( 'Confidence in assessment 0-1' ),
          issues: z.array( z.string() ).optional().describe( 'Any factual issues found' )
        } )
      } )
    } );

    return new EvaluationBooleanResult( {
      value: output.isFactual,
      confidence: output.confidence,
      reasoning: output.issues?.length ?
        `Issues found: ${output.issues.join( ', ' )}` :
        'No factual issues detected'
    } );
  }
} );
LLM String Evaluation - Content Classification
typescript
import { evaluator, z, EvaluationStringResult } from '@outputai/core';
import { generateText, aiSdk } from '@outputai/llm';

export const evaluateContentCategory = evaluator( {
  name: 'evaluate_content_category',
  description: 'Classify content into a category',
  inputSchema: z.object( {
    content: z.string(),
    categories: z.array( z.string() )
  } ),
  fn: async ( { content, categories } ) => {
    const { output } = await generateText( {
      prompt: 'categorize_content@v1',
      variables: {
        content,
        categories: categories.join( ', ' )
      },
      output: aiSdk.Output.object( {
        schema: z.object( {
          category: z.string().describe( 'The best matching category' ),
          confidence: z.number().describe( 'Confidence in classification 0-1' ),
          explanation: z.string().describe( 'Why this category was chosen' )
        } )
      } )
    } );

    return new EvaluationStringResult( {
      value: output.category,
      confidence: output.confidence,
      reasoning: output.explanation
    } );
  }
} );
Show full SKILL.md (255 more words)Show less

EvaluationResult with Feedback

Use the feedback field to provide actionable improvement suggestions alongside your evaluation result. Import EvaluationFeedback from @outputai/core to create feedback objects.

typescript
import { evaluator, z, EvaluationStringResult, EvaluationFeedback } from '@outputai/core';

export const evaluateWithFeedback = evaluator( {
  name: 'evaluate_with_feedback',
  description: 'Evaluate content quality and provide actionable feedback',
  inputSchema: z.string(),
  fn: async response => {
    const feedback = [];

    if ( response.length < 50 ) {
      feedback.push( new EvaluationFeedback( {
        issue: 'Response is too short',
        suggestion: 'Expand the response with more detail',
        priority: 'medium'
      } ) );
    }

    return new EvaluationStringResult( {
      value: feedback.length === 0 ? 'good' : 'needs_improvement',
      confidence: 0.85,
      feedback: feedback
    } );
  }
} );
EvaluationFeedback Properties
PropertyTypeDescription
issuestringThe problem identified
suggestionstringRecommended fix
prioritystringPriority level (e.g., 'low', 'medium', 'high')

Multi-Dimensional Evaluation

Use the dimensions field to nest EvaluationResult instances for sub-scores. Each dimension should use the name field to identify it.

typescript
import { evaluator, z, EvaluationStringResult, EvaluationNumberResult } from '@outputai/core';

export const evaluateMultiDimensional = evaluator( {
  name: 'evaluate_multi_dimensional',
  description: 'Evaluate content across multiple quality dimensions',
  inputSchema: z.string(),
  fn: async response => {
    const coherenceScore = calculateCoherence( response );
    const relevanceScore = calculateRelevance( response );
    const overallScore = ( coherenceScore + relevanceScore ) / 2;

    return new EvaluationStringResult( {
      value: overallScore > 0.7 ? 'high_quality' : 'low_quality',
      confidence: 0.9,
      dimensions: [
        new EvaluationNumberResult( {
          value: coherenceScore,
          confidence: 0.85,
          name: 'coherence'
        } ),
        new EvaluationNumberResult( {
          value: relevanceScore,
          confidence: 0.88,
          name: 'relevance'
        } )
      ]
    } );
  }
} );

Complete Example

Based on a real workflow evaluator file:

typescript
import { evaluator, z, EvaluationBooleanResult, EvaluationNumberResult } from '@outputai/core';
import { generateText, aiSdk } from '@outputai/llm';
import { blogContentSchema } from './types.js';
import type { BlogContent, QualityMetrics } from './types.js';

// Simple boolean evaluator
export const evaluateMinimumLength = evaluator( {
  name: 'evaluate_minimum_length',
  description: 'Check if blog content meets minimum length requirements',
  inputSchema: blogContentSchema,
  fn: async ( input: BlogContent ) => {
    const MIN_TOKENS = 500;
    const meetsRequirement = input.tokenCount >= MIN_TOKENS;

    return new EvaluationBooleanResult( {
      value: meetsRequirement,
      confidence: 1.0,
      reasoning: `Content has ${input.tokenCount} tokens (minimum: ${MIN_TOKENS})`
    } );
  }
} );

// LLM-powered number evaluator
export const evaluateSignalToNoise = evaluator( {
  name: 'evaluate_signal_to_noise',
  description: 'Evaluate the signal-to-noise ratio of blog content',
  inputSchema: blogContentSchema,
  fn: async ( input: BlogContent ) => {
    const { output } = await generateText( {
      prompt: 'signal_noise@v1',
      variables: {
        title: input.title,
        content: input.content
      },
      output: aiSdk.Output.object( {
        schema: z.object( {
          score: z.number().describe( 'Signal-to-noise score 0-100' )
        } )
      } )
    } );

    return new EvaluationNumberResult( {
      value: output.score,
      confidence: 0.85
    } );
  }
} );

// LLM-powered boolean evaluator
export const evaluateRelevance = evaluator( {
  name: 'evaluate_relevance',
  description: 'Check if content is relevant to the stated topic',
  inputSchema: z.object( {
    content: z.string(),
    topic: z.string(),
    keywords: z.array( z.string() )
  } ),
  fn: async ( { content, topic, keywords } ) => {
    const { output } = await generateText( {
      prompt: 'relevance_check@v1',
      variables: { content, topic, keywords: keywords.join( ', ' ) },
      output: aiSdk.Output.object( {
        schema: z.object( {
          isRelevant: z.boolean(),
          relevanceScore: z.number().describe( 'Relevance score 0-1' ),
          explanation: z.string()
        } )
      } )
    } );

    return new EvaluationBooleanResult( {
      value: output.isRelevant,
      confidence: output.relevanceScore,
      reasoning: output.explanation
    } );
  }
} );

Best Practices

1. Use Appropriate Result Types
typescript
// Boolean for pass/fail decisions
return new EvaluationBooleanResult( { value: true, confidence: 0.9 } );

// Number for scores and ratings
return new EvaluationNumberResult( { value: 85, confidence: 0.85 } );

// String for categories, labels, or classifications
return new EvaluationStringResult( { value: 'positive', confidence: 0.9 } );
2. Provide Meaningful Confidence Scores
typescript
// High confidence for deterministic checks
confidence: 1.0  // e.g., length checks, pattern matching

// Medium confidence for heuristic-based evaluations
confidence: 0.85  // e.g., LLM-based assessments

// Lower confidence for uncertain evaluations
confidence: 0.7  // e.g., subjective quality judgments
3. Include Reasoning for Transparency
typescript
return new EvaluationBooleanResult( {
  value: false,
  confidence: 0.95,
  reasoning: `Content contains ${errorCount} grammatical errors, exceeding threshold of ${maxErrors}`
} );
4. Keep Evaluators Focused
typescript
// Good - single responsibility
export const evaluateGrammar = evaluator( { ... } );
export const evaluateReadability = evaluator( { ... } );
export const evaluateTone = evaluator( { ... } );

// Avoid - doing too much in one evaluator
export const evaluateEverything = evaluator( { ... } );
5. Use Descriptive Names
typescript
// Good - clear what is being evaluated
name: 'evaluate_content_originality'
name: 'evaluate_factual_accuracy'
name: 'evaluate_sentiment_alignment'

// Avoid - vague names
name: 'check'
name: 'validate'
name: 'evaluate_stuff'
6. Use Feedback for Actionable Improvements
typescript
feedback: [
  new EvaluationFeedback( {
    issue: 'Missing conclusion paragraph',
    suggestion: 'Add a summary paragraph at the end',
    priority: 'high'
  } )
]
7. Use Dimensions for Multi-Criteria Evaluation
typescript
dimensions: [
  new EvaluationNumberResult( { value: 8, confidence: 0.9, name: 'coherence' } ),
  new EvaluationNumberResult( { value: 6, confidence: 0.85, name: 'relevance' } )
]

Verification Checklist

  • evaluator, z, result types imported from @outputai/core
  • generateText and aiSdk imported from @outputai/llm if using LLM (not direct provider)
  • LLM output schemas use .describe() instead of .min()/.max() on z.number()
  • All imports use .js extension
  • Named exports used for each evaluator
  • Each evaluator has name, description, inputSchema, fn
  • Evaluator name uses snake_case
  • Returns appropriate result type (EvaluationBooleanResult, EvaluationNumberResult, or EvaluationStringResult)
  • Confidence score between 0.0 and 1.0
  • Evaluators only import allowed dependencies (local files, shared code)
  • No imports of other evaluators, steps, or workflows
  • EvaluationFeedback imported from @outputai/core when using feedback
  • Feedback objects include issue, suggestion, and priority
  • Dimensions use the name field to identify sub-evaluations
  • output-dev-workflow-function - Orchestrating evaluators in workflow.ts
  • output-dev-step-function - Creating step functions
  • output-dev-types-file - Defining evaluator input schemas
  • output-dev-prompt-file - Creating prompt files for LLM-powered evaluators
  • output-dev-folder-structure - Understanding project layout
  • output-eval-error-analysis — Identify what to evaluate before writing evaluators

© growthxai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function of growthxai/output.

Open the folder on GitHubat commit 52b51ac

Compare with similar skills

Output Dev Evaluator Function next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Output Dev Evaluator Function compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Output Dev Evaluator Function this skillgrowthxai/output440—~5.5kAutomated safety check: PassApache-2.0
Arize Evaluatorgithub/awesome-copilot40k2 repos~8.1kAutomated safety check: NotesMIT
LLM Evaluationdavila7/claude-code-templates32k13 repos~3.5kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT
EvaluatorsArize-ai/phoenix12k—~1.7kAutomated safety check: PassCustom licence
Agent Evaluation Reportingsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 2 repos~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 13 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Evaluators

    Arize-ai/phoenix

    Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

    12k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check passed
  • Agent Evaluation Reporting

    sickn33/agentic-awesome-skills

    A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Agent WorkflowsAuto-check passed
  • Official

    Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified.

    40k GitHub stars~6.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from growthxai/output

All 52 skills in this repo
  • Zod schema constraints that Anthropic rejects or silently ignores when sent as structured-output tool definitions via aiSdk.Output.object().

    440 GitHub stars~597 tokensUpdated today
    Auto-check passed
  • Output Build Workflow

    growthxai/output

    Implement an Output SDK workflow from a plan document. An agent skill from growthxai/output.

    440 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Output Credentials Edit

    growthxai/output

    View, edit, and set encrypted credentials in an Output.ai project.

    440 GitHub stars~1.1k tokensUpdated today
    Auto-check: notes
  • Wire encrypted credentials to environment variables using the credential: convention.

    440 GitHub stars~930 tokensUpdated today
    Auto-check: notes
  • Output Credentials Init

    growthxai/output

    Initialize encrypted credentials for an Output.ai project. An agent skill from growthxai/output.

    440 GitHub stars~803 tokensUpdated today
    Auto-check: notes
  • Output Debug Workflow

    growthxai/output

    Debug Output SDK workflow issues. An agent skill from growthxai/output.

    440 GitHub stars~1.5k tokensUpdated today
    Auto-check passed

Questions about Output Dev Evaluator Function

What does Output Dev Evaluator Function do?

Create evaluator functions in evaluators.ts for Output SDK workflows. Output Dev Evaluator Function is an agent skill from growthxai/output.ts for Output SDK workflows.

When should I use Output Dev Evaluator Function?

Output Dev Evaluator Function fits situations like: implementing quality assessment; validation logic; content evaluation.

How do I install Output Dev Evaluator Function in Claude Code?

Run `npx skills add growthxai/output --skill output-dev-evaluator-function -a claude-code`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function in growthxai/output) into .claude/skills/output-dev-evaluator-function in your project. Claude Code loads it when a task matches its description.

How do I install Output Dev Evaluator Function in Codex?

Run `npx skills add growthxai/output --skill output-dev-evaluator-function -a codex`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-dev-evaluator-function in growthxai/output) into .agents/skills/output-dev-evaluator-function in your project. Codex loads it when a task matches its description.

Can I use Output Dev Evaluator Function in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add growthxai/output --skill output-dev-evaluator-function -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/output-dev-evaluator-function, .gemini/skills/output-dev-evaluator-function, .github/skills/output-dev-evaluator-function and .opencode/skills/output-dev-evaluator-function in your project.

What does Output Dev Evaluator Function need to run?

SKILL.md names no scripts, command-line tools or credentials: Output Dev Evaluator Function is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Write, Edit.

Does Output Dev Evaluator Function access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Output Dev Evaluator Function safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Output Dev Evaluator Function use?

Output Dev Evaluator Function is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Output Dev Evaluator Function use?

About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Output Dev Evaluator Function?

Skills that share tags, products or a category with Output Dev Evaluator Function: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars) and Evaluators (Arize-ai/phoenix, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Output Dev Evaluator Function?

growthxai (a GitHub organization) maintains it in growthxai/output, which has 440 GitHub stars. The repository holds 52 skills in this directory. The repository was last updated on October 7, 2026.

Source: growthxai/output on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.