Agent skill

Output Dev Eval Testing

by growthxai in growthxai/output

Create offline evaluation tests for Output SDK workflows using @outputai/evals.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Output Dev Eval Testing

skills CLI
$ npx skills add growthxai/output --skill output-dev-eval-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install growthxai/output output-dev-eval-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .claude/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing .claude/skills/output-dev-eval-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
output-dev-eval-testing
GitHub stars
440
Token cost
~3.8k tokens
SKILL.md length
940 words
Files
1
Skills in repo
52
Repo updated
First seen
Licence
Apache-2.0

At a glance

Create offline evaluation tests for Output SDK workflows using @outputai/evals.

  • Works in 5 steps: Loads all dataset YAML files from… → Without --cached: executes the workflow… → Sends all datasets to the… → …
  • Implementing test evaluators with verify()
  • SKILL.md covers Overview, When to Use This Skill, Directory Structure and Creating Evaluators with…, plus 7 more sections
  • Calls npm; reaches stripe.com

What it does

Output Dev Eval Testing is an agent skill from growthxai/output. Create offline evaluation tests for Output SDK workflows using @outputai/evals. Use when implementing test evaluators with verify(), creating dataset YAML files, building eval workflows, or running workflow tests via CLI.

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already… The licence is Apache-2.0.

When your agent uses it

  • Implementing test evaluators with verify()
  • Creating dataset YAML files
  • Building eval workflows
  • Running workflow tests via CLI

Example prompts

  • “/output-dev-eval-testing”

Requirements

  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Loads all dataset YAML files from tests/datasets/
  2. Without --cached: executes the workflow for each dataset to get fresh output
  3. Sends all datasets to the {workflow_name}_eval workflow
  4. Reports per-dataset and per-evaluator verdicts
  5. Exits with code 1 if any required evaluator fails

What it can do on your machine

Read from SKILL.md and the folder at commit 52b51ac. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • stripe.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Output Dev Eval Testing loads about 3.8k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 940 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from growthxai/output at commit 52b51ac, republished under its Apache-2.0 licence (© growthxai). 940 words, ~3,833 tokens.

Download SKILL.mdSave it as .claude/skills/output-dev-eval-testing/SKILL.md (or your agent's skills folder).
name
output-dev-eval-testing
description
Create offline evaluation tests for Output SDK workflows using @outputai/evals. Use when implementing test evaluators with verify(), creating dataset YAML files, building eval workflows, or running workflow tests via CLI.
allowed-tools
Bash, Read, Write, Edit

Offline Evaluation Testing

Overview

The @outputai/evals package provides an offline evaluation framework for testing workflow quality using datasets and evaluators. This is complementary to the runtime evaluator() from @outputai/core:

AspectRuntime Evaluators (@outputai/core)Offline Eval Tests (@outputai/evals)
WhenDuring workflow executionAfter execution, at test time
Whereevaluators.ts in workflow foldertests/evals/ in workflow folder
PurposeLive quality scoring with confidenceDataset-driven pass/fail verification
Triggered byWorkflow orchestrationoutput workflow test CLI command
ReturnsEvaluationBooleanResult, etc.Verdict helpers (pass/partial/fail)

Use offline eval testing when you want to validate workflow behavior against known datasets, build regression test suites, or assess subjective quality with LLM judges.

When to Use This Skill

  • Creating files in tests/evals/ or tests/datasets/
  • Writing evaluators that use verify() from @outputai/evals
  • Creating YAML dataset files for test cases
  • Building eval workflows with evalWorkflow()
  • Running output workflow test commands
  • Setting up ground truth data for evaluators

Directory Structure

Add a tests/ directory inside the workflow folder:

src/workflows/{workflow_name}/
├── workflow.ts
├── steps.ts
├── evaluators.ts          # Runtime evaluators (optional)
├── types.ts
└── tests/
    ├── datasets/
    │   ├── happy_path.yml
    │   └── edge_case.yml
    └── evals/
        ├── evaluators.ts  # Offline eval test evaluators
        ├── workflow.ts     # Eval workflow definition
        └── judge_topic@v1.prompt  # LLM judge prompts (optional)

Creating Evaluators with verify()

Import verify and Verdict from @outputai/evals (not @outputai/core):

typescript
// tests/evals/evaluators.ts
import { verify, Verdict } from '@outputai/evals';
import { z } from '@outputai/core';
verify() Signature
typescript
verify(options, checkFn)

Options:

  • name — unique evaluator identifier (snake_case)
  • input — Zod schema for the workflow input (optional, defaults to z.any())
  • output — Zod schema for the workflow output (optional, defaults to z.any())

Check function receives:

typescript
{
  input,    // typed workflow input
  output,   // typed workflow output
  context: {
    ground_truth: Record<string, unknown>  // from dataset YAML
  }
}

Returns: any Verdict helper result.

Basic Example
typescript
import { verify, Verdict } from '@outputai/evals';
import { z } from '@outputai/core';

export const evaluateSum = verify(
  {
    name: 'evaluate_sum',
    input: z.object({ values: z.array(z.number()) }),
    output: z.object({ result: z.number() })
  },
  ({ input, output }) =>
    Verdict.equals(output.result, input.values.reduce((a, b) => a + b, 0))
);
Using Ground Truth

Ground truth values come from the dataset YAML and are available via context.ground_truth:

typescript
export const lengthCheck = verify(
  { name: 'length_check', input: blogInput, output: blogOutput },
  ({ output, context }) =>
    Verdict.gte(output.blog_post.length, Number(context.ground_truth.min_length ?? 100))
);

Verdict Helpers

All deterministic helpers return results with confidence 1.0.

Equality & Comparison
MethodDescription
Verdict.equals(actual, expected)Strict equality (===)
Verdict.closeTo(actual, expected, tolerance)Within numeric tolerance
Verdict.gt(actual, threshold)Greater than
Verdict.gte(actual, threshold)Greater than or equal
Verdict.lt(actual, threshold)Less than
Verdict.lte(actual, threshold)Less than or equal
Verdict.inRange(actual, min, max)Within inclusive range
String & Array
MethodDescription
Verdict.contains(haystack, needle)String includes substring
Verdict.matches(value, pattern)Regex match
Verdict.includesAll(actual, expected)Array contains all expected values
Verdict.includesAny(actual, expected)Array contains at least one expected value
Boolean
MethodDescription
Verdict.isTrue(value)Value is true
Verdict.isFalse(value)Value is false
Manual Verdicts
MethodDescription
Verdict.pass(reasoning?)Explicit pass
Verdict.partial(confidence, reasoning?, feedback?)Partial pass with confidence
Verdict.fail(reasoning, feedback?)Explicit fail

LLM Judge Evaluators

Before writing a judge prompt, identify the specific failure mode via error analysis (output-eval-error-analysis). Design the judge following output-eval-judge-prompt. After writing it, validate against human labels using output-eval-validate-judge.

For subjective quality assessments, use judge functions with .prompt files:

typescript
import { verify, judgeVerdict, judgeScore, judgeLabel } from '@outputai/evals';

// Returns pass/partial/fail verdict from an LLM
export const evaluateTopic = verify(
  { name: 'evaluate_topic', input: blogInput, output: blogOutput },
  async ({ input, output, context }) =>
    judgeVerdict({
      prompt: 'judge_topic@v1',
      variables: {
        blog_title: output.title,
        blog_post: output.blog_post,
        required_topic: String(context.ground_truth.required_topic ?? input.topic)
      }
    })
);

// Returns a numeric score from an LLM
export const evaluateQuality = verify(
  { name: 'evaluate_quality', input: blogInput, output: blogOutput },
  async ({ input, output }) =>
    judgeScore({
      prompt: 'judge_quality@v1',
      variables: { blog_title: output.title, blog_post: output.blog_post, topic: input.topic }
    })
);

// Returns a string label from an LLM
export const evaluateTone = verify(
  { name: 'evaluate_tone', input: blogInput, output: blogOutput },
  async ({ output }) =>
    judgeLabel({
      prompt: 'judge_tone@v1',
      variables: { blog_title: output.title, blog_post: output.blog_post }
    })
);
Judge .prompt File Format

Judge prompt files live alongside evaluators in tests/evals/:

yaml
# tests/evals/judge_topic@v1.prompt
---
provider: anthropic
# current as of 2026-05-04 — run output-dev-model-selection for the latest
model: claude-haiku-4-5-20251001
temperature: 0
maxOutputTokens: 1000
---

<system>
You are an evaluation judge. Assess whether a blog post is faithfully about the required topic.

Return a JSON object with:
- verdict: "pass" if the blog clearly focuses on the topic, "partial" if it mentions the topic but lacks depth, "fail" if it is not about the topic
- reasoning: a brief explanation of your judgment
</system>

<user>
Required topic: {{ required_topic }}

Blog title: {{ blog_title }}

Blog post:
{{ blog_post }}

Judge whether this blog post is faithfully about the required topic.
</user>

Creating Eval Workflows

The eval workflow wires evaluators together and defines how to interpret results.

typescript
// tests/evals/workflow.ts
import { evalWorkflow } from '@outputai/evals';
import { evaluateSum } from './evaluators.js';

export default evalWorkflow({
  name: 'simple_eval',
  evals: [
    {
      evaluator: evaluateSum,
      criticality: 'required',
      interpret: { type: 'boolean' }
    }
  ]
});
Eval Definition Fields

Each entry in the evals array has:

  • evaluator — the function created by verify()
  • criticality — 'required' (affects pass/fail) or 'informational' (reported but doesn't block)
  • interpret — how to convert the evaluator's return value into a verdict
Interpret Types
TypeEvaluator ReturnsMapping
{ type: 'boolean' }Verdict.equals(), Verdict.gte(), etc.true = pass, false = fail
{ type: 'verdict' }judgeVerdict() or Verdict.pass/partial/fail()Direct pass-through
{ type: 'number', pass: 0.7, partial: 0.4 }judgeScore()>=pass = pass, >=partial = partial, else fail
{ type: 'string', pass: ['a', 'b'], partial: ['c'] }judgeLabel()Label in pass list = pass, in partial list = partial, else fail
Full Example with Mixed Evaluators
typescript
export default evalWorkflow({
  name: 'blog_generator_eval',
  evals: [
    {
      evaluator: lengthOfOutput,
      criticality: 'required',
      interpret: { type: 'boolean' }
    },
    {
      evaluator: evaluateTopic,
      criticality: 'required',
      interpret: { type: 'verdict' }
    },
    {
      evaluator: evaluateQuality,
      criticality: 'required',
      interpret: { type: 'number', pass: 0.7, partial: 0.4 }
    },
    {
      evaluator: evaluateContent,
      criticality: 'informational',
      interpret: { type: 'boolean' }
    },
    {
      evaluator: evaluateTone,
      criticality: 'informational',
      interpret: { type: 'string', pass: ['professional', 'informative'], partial: ['casual'] }
    }
  ]
});
Naming Convention

The eval workflow name must end in _eval and match the pattern {workflow_name}_eval. The CLI resolves this automatically — output workflow test blog_generator looks for blog_generator_eval.

Dataset Files

For methodology on designing diverse datasets that cover failure-prone regions, see output-eval-dataset-design.

Datasets are YAML files in tests/datasets/. Each file represents one test case.

Basic Format
yaml
name: basic_input
input:
  values:
    - 1
    - 2
    - 3
    - 4
    - 5
last_output:
  output:
    result: 15
  executionTimeMs: 100
  date: '2026-02-13T00:00:00.000Z'
Show full SKILL.md (389 more words)Show less
With Ground Truth

Ground truth provides expected values for evaluators. You can set global values and per-evaluator overrides:

yaml
name: stripe_blog
input:
  topic: "Stripe the payment processor"
  requirements: "Include a link to https://stripe.com/en-gb/pricing"
last_output:
  output:
    title: "Stripe: The Modern Payment Processing Platform"
    blog_post: |
      Stripe has revolutionized online payment processing...
  executionTimeMs: 5000
  date: '2026-02-16T00:00:00.000Z'
ground_truth:
  notes: "Known good case"
  evals:
    length_of_output:
      min_length: 100
    evaluate_topic:
      required_topic: "Stripe the payment processor"
    evaluate_content:
      required_content: "https://stripe.com/en-gb/pricing"

The ground_truth.evals.<evaluator_name> values are merged with the top-level ground truth and passed to the evaluator via context.ground_truth.

CLI Commands

output workflow test <workflow_name>

Runs evaluations against all datasets for a workflow.

FlagDescription
--cachedUse cached output from dataset files (skip workflow execution)
--saveRun workflow fresh and save output + eval results back to dataset files
--dataset <names>Comma-separated list of dataset names to run (default: all)
--jsonOutput machine-readable JSON instead of the rendered report

Execution flow:

  1. Loads all dataset YAML files from tests/datasets/
  2. Without --cached: executes the workflow for each dataset to get fresh output
  3. Sends all datasets to the {workflow_name}_eval workflow
  4. Reports per-dataset and per-evaluator verdicts
  5. Exits with code 1 if any required evaluator fails
output workflow dataset list <workflow_name>

Lists all datasets for a workflow with their cached status.

FlagDescription
--format <type>Output format: table (default) or text
--jsonOutput machine-readable JSON
output workflow dataset generate <workflow_name> [scenario]

Generates a new dataset file by running the workflow.

FlagDescription
--input <json>Workflow input as a JSON string or file path
--name <name>Dataset filename (defaults to scenario name)
--trace <path>Generate from a local trace file instead of running the workflow
--downloadDownload traces from S3 and convert to datasets
--limit <n>Max traces to download from S3 (default: 5)
Common Usage
bash
# Generate dataset from inline JSON input
output workflow dataset generate my_workflow --input '{"key": "value"}' --name my_test

# Generate from a scenario file
output workflow dataset generate my_workflow basic

# Run evals with cached output (fast, no re-execution)
output workflow test my_workflow --cached

# Run evals fresh and save results
output workflow test my_workflow --save

# Run specific datasets only
output workflow test my_workflow --dataset happy_path,edge_case

# List all datasets
output workflow dataset list my_workflow

Typical Workflow

bash
# 1. Start the dev server
npm run output:dev

# 2. Generate datasets from real workflow runs
output workflow dataset generate blog_generator --input '{"topic": "AI"}' --name ai_post

# 3. Edit the dataset YAML to add ground_truth values for your evaluators

# 4. Run evals with --save to cache output and eval results
output workflow test blog_generator --save

# 5. Iterate on evaluators, re-run with cached output (fast)
output workflow test blog_generator --cached

# 6. List all datasets
output workflow dataset list blog_generator

Verification Checklist

  • Evaluators import verify, Verdict from @outputai/evals (not @outputai/core)
  • Eval workflow imports evalWorkflow from @outputai/evals
  • All imports use .js extension
  • Eval workflow name follows {workflow_name}_eval pattern
  • Dataset YAML files are in tests/datasets/
  • Evaluator files are in tests/evals/
  • Each evaluator has a unique name in snake_case
  • criticality is set to 'required' or 'informational' for each eval
  • interpret type matches evaluator return type
  • Ground truth keys in dataset match evaluator names
  • Judge .prompt files are in tests/evals/ alongside evaluators
  • z is imported from @outputai/core (not zod)
  • output-dev-evaluator-function — Runtime evaluators using evaluator() from @outputai/core
  • output-dev-scenario-file — Creating scenario JSON files for workflow execution
  • output-dev-folder-structure — Understanding project directory layout
  • output-dev-prompt-file — Creating .prompt files for LLM operations
  • output-eval-error-analysis — Identify failure modes before building evaluators
  • output-eval-judge-prompt — Design effective LLM judge prompts
  • output-eval-dataset-design — Generate diverse test datasets
  • output-eval-validate-judge — Validate LLM judges against human labels
  • output-eval-audit — Audit an existing eval suite for trustworthiness

© growthxai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing of growthxai/output.

Open the folder on GitHubat commit 52b51ac

Compare with similar skills

Output Dev Eval Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Output Dev Eval Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Output Dev Eval Testing this skillgrowthxai/output440—~3.8kAutomated safety check: NotesApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from growthxai/output

All 52 skills in this repo
  • Zod schema constraints that Anthropic rejects or silently ignores when sent as structured-output tool definitions via aiSdk.Output.object().

    440 GitHub stars~597 tokensUpdated yesterday
    Auto-check passed
  • Output Build Workflow

    growthxai/output

    Implement an Output SDK workflow from a plan document. An agent skill from growthxai/output.

    440 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Output Credentials Edit

    growthxai/output

    View, edit, and set encrypted credentials in an Output.ai project.

    440 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check: notes
  • Wire encrypted credentials to environment variables using the credential: convention.

    440 GitHub stars~930 tokensUpdated yesterday
    Auto-check: notes
  • Output Credentials Init

    growthxai/output

    Initialize encrypted credentials for an Output.ai project. An agent skill from growthxai/output.

    440 GitHub stars~803 tokensUpdated yesterday
    Auto-check: notes
  • Output Debug Workflow

    growthxai/output

    Debug Output SDK workflow issues. An agent skill from growthxai/output.

    440 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Questions about Output Dev Eval Testing

What does Output Dev Eval Testing do?

Create offline evaluation tests for Output SDK workflows using @outputai/evals. Output Dev Eval Testing is an agent skill from growthxai/output. Create offline evaluation tests for Output SDK workflows using @outputai/evals.

When should I use Output Dev Eval Testing?

Output Dev Eval Testing fits situations like: implementing test evaluators with verify(); creating dataset YAML files; building eval workflows; running workflow tests via CLI.

How do I install Output Dev Eval Testing in Claude Code?

Run `npx skills add growthxai/output --skill output-dev-eval-testing -a claude-code`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing in growthxai/output) into .claude/skills/output-dev-eval-testing in your project. Claude Code loads it when a task matches its description.

How do I install Output Dev Eval Testing in Codex?

Run `npx skills add growthxai/output --skill output-dev-eval-testing -a codex`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing in growthxai/output) into .agents/skills/output-dev-eval-testing in your project. Codex loads it when a task matches its description.

Can I use Output Dev Eval Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add growthxai/output --skill output-dev-eval-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/output-dev-eval-testing, .gemini/skills/output-dev-eval-testing, .github/skills/output-dev-eval-testing and .opencode/skills/output-dev-eval-testing in your project.

What does Output Dev Eval Testing need to run?

Going by SKILL.md and its folder, Output Dev Eval Testing needs the command-line tools its instructions call (npm). Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.

Does Output Dev Eval Testing access the network?

SKILL.md names 1 domain. In commands or code: stripe.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Output Dev Eval Testing safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Output Dev Eval Testing use?

Output Dev Eval Testing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Output Dev Eval Testing use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Output Dev Eval Testing?

Skills that share tags, products or a category with Output Dev Eval Testing: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Output Dev Eval Testing?

growthxai (a GitHub organization) maintains it in growthxai/output, which has 440 GitHub stars. The repository holds 52 skills in this directory. The repository was last updated on October 7, 2026.

Source: growthxai/output on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.