LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Create offline evaluation tests for Output SDK workflows using @outputai/evals.
$ npx skills add growthxai/output --skill output-dev-eval-testing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install growthxai/output output-dev-eval-testing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .claude/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing .claude/skills/output-dev-eval-testing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "output-dev-eval-testing" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing into .claude/skills/output-dev-eval-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-eval-testing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add growthxai/output --skill output-dev-eval-testing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install growthxai/output output-dev-eval-testing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .agents/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing .agents/skills/output-dev-eval-testing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "output-dev-eval-testing" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing into .agents/skills/output-dev-eval-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-eval-testing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add growthxai/output --skill output-dev-eval-testing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install growthxai/output output-dev-eval-testing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing .cursor/skills/output-dev-eval-testing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "output-dev-eval-testing" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing into .cursor/skills/output-dev-eval-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-eval-testing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/growthxai/output.git --path coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add growthxai/output --skill output-dev-eval-testing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install growthxai/output output-dev-eval-testing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing .gemini/skills/output-dev-eval-testing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "output-dev-eval-testing" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing into .gemini/skills/output-dev-eval-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-eval-testing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install growthxai/output output-dev-eval-testingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add growthxai/output --skill output-dev-eval-testing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .github/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing .github/skills/output-dev-eval-testing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "output-dev-eval-testing" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing into .github/skills/output-dev-eval-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-eval-testing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add growthxai/output --skill output-dev-eval-testing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install growthxai/output output-dev-eval-testing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing .opencode/skills/output-dev-eval-testing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "output-dev-eval-testing" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing into .opencode/skills/output-dev-eval-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-dev-eval-testing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
output-dev-eval-testingCreate offline evaluation tests for Output SDK workflows using @outputai/evals.
Output Dev Eval Testing is an agent skill from growthxai/output. Create offline evaluation tests for Output SDK workflows using @outputai/evals. Use when implementing test evaluators with verify(), creating dataset YAML files, building eval workflows, or running workflow tests via CLI.
Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already… The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 52b51ac. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadWriteEditFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npmFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
stripe.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Output Dev Eval Testing loads about 3.8k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 940 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, Read, Write, EditAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from growthxai/output at commit 52b51ac, republished under its Apache-2.0 licence (© growthxai). 940 words, ~3,833 tokens.
.claude/skills/output-dev-eval-testing/SKILL.md (or your agent's skills folder).The @outputai/evals package provides an offline evaluation framework for testing workflow quality using datasets and evaluators. This is complementary to the runtime evaluator() from @outputai/core:
| Aspect | Runtime Evaluators (@outputai/core) | Offline Eval Tests (@outputai/evals) |
|---|---|---|
| When | During workflow execution | After execution, at test time |
| Where | evaluators.ts in workflow folder | tests/evals/ in workflow folder |
| Purpose | Live quality scoring with confidence | Dataset-driven pass/fail verification |
| Triggered by | Workflow orchestration | output workflow test CLI command |
| Returns | EvaluationBooleanResult, etc. | Verdict helpers (pass/partial/fail) |
Use offline eval testing when you want to validate workflow behavior against known datasets, build regression test suites, or assess subjective quality with LLM judges.
tests/evals/ or tests/datasets/verify() from @outputai/evalsevalWorkflow()output workflow test commandsAdd a tests/ directory inside the workflow folder:
src/workflows/{workflow_name}/
├── workflow.ts
├── steps.ts
├── evaluators.ts # Runtime evaluators (optional)
├── types.ts
└── tests/
├── datasets/
│ ├── happy_path.yml
│ └── edge_case.yml
└── evals/
├── evaluators.ts # Offline eval test evaluators
├── workflow.ts # Eval workflow definition
└── judge_topic@v1.prompt # LLM judge prompts (optional)verify()Import verify and Verdict from @outputai/evals (not @outputai/core):
// tests/evals/evaluators.ts
import { verify, Verdict } from '@outputai/evals';
import { z } from '@outputai/core';verify() Signatureverify(options, checkFn)Options:
name — unique evaluator identifier (snake_case)input — Zod schema for the workflow input (optional, defaults to z.any())output — Zod schema for the workflow output (optional, defaults to z.any())Check function receives:
{
input, // typed workflow input
output, // typed workflow output
context: {
ground_truth: Record<string, unknown> // from dataset YAML
}
}Returns: any Verdict helper result.
import { verify, Verdict } from '@outputai/evals';
import { z } from '@outputai/core';
export const evaluateSum = verify(
{
name: 'evaluate_sum',
input: z.object({ values: z.array(z.number()) }),
output: z.object({ result: z.number() })
},
({ input, output }) =>
Verdict.equals(output.result, input.values.reduce((a, b) => a + b, 0))
);Ground truth values come from the dataset YAML and are available via context.ground_truth:
export const lengthCheck = verify(
{ name: 'length_check', input: blogInput, output: blogOutput },
({ output, context }) =>
Verdict.gte(output.blog_post.length, Number(context.ground_truth.min_length ?? 100))
);All deterministic helpers return results with confidence 1.0.
| Method | Description |
|---|---|
Verdict.equals(actual, expected) | Strict equality (===) |
Verdict.closeTo(actual, expected, tolerance) | Within numeric tolerance |
Verdict.gt(actual, threshold) | Greater than |
Verdict.gte(actual, threshold) | Greater than or equal |
Verdict.lt(actual, threshold) | Less than |
Verdict.lte(actual, threshold) | Less than or equal |
Verdict.inRange(actual, min, max) | Within inclusive range |
| Method | Description |
|---|---|
Verdict.contains(haystack, needle) | String includes substring |
Verdict.matches(value, pattern) | Regex match |
Verdict.includesAll(actual, expected) | Array contains all expected values |
Verdict.includesAny(actual, expected) | Array contains at least one expected value |
| Method | Description |
|---|---|
Verdict.isTrue(value) | Value is true |
Verdict.isFalse(value) | Value is false |
| Method | Description |
|---|---|
Verdict.pass(reasoning?) | Explicit pass |
Verdict.partial(confidence, reasoning?, feedback?) | Partial pass with confidence |
Verdict.fail(reasoning, feedback?) | Explicit fail |
Before writing a judge prompt, identify the specific failure mode via error analysis (output-eval-error-analysis). Design the judge following output-eval-judge-prompt. After writing it, validate against human labels using output-eval-validate-judge.
For subjective quality assessments, use judge functions with .prompt files:
import { verify, judgeVerdict, judgeScore, judgeLabel } from '@outputai/evals';
// Returns pass/partial/fail verdict from an LLM
export const evaluateTopic = verify(
{ name: 'evaluate_topic', input: blogInput, output: blogOutput },
async ({ input, output, context }) =>
judgeVerdict({
prompt: 'judge_topic@v1',
variables: {
blog_title: output.title,
blog_post: output.blog_post,
required_topic: String(context.ground_truth.required_topic ?? input.topic)
}
})
);
// Returns a numeric score from an LLM
export const evaluateQuality = verify(
{ name: 'evaluate_quality', input: blogInput, output: blogOutput },
async ({ input, output }) =>
judgeScore({
prompt: 'judge_quality@v1',
variables: { blog_title: output.title, blog_post: output.blog_post, topic: input.topic }
})
);
// Returns a string label from an LLM
export const evaluateTone = verify(
{ name: 'evaluate_tone', input: blogInput, output: blogOutput },
async ({ output }) =>
judgeLabel({
prompt: 'judge_tone@v1',
variables: { blog_title: output.title, blog_post: output.blog_post }
})
);.prompt File FormatJudge prompt files live alongside evaluators in tests/evals/:
# tests/evals/judge_topic@v1.prompt
---
provider: anthropic
# current as of 2026-05-04 — run output-dev-model-selection for the latest
model: claude-haiku-4-5-20251001
temperature: 0
maxOutputTokens: 1000
---
<system>
You are an evaluation judge. Assess whether a blog post is faithfully about the required topic.
Return a JSON object with:
- verdict: "pass" if the blog clearly focuses on the topic, "partial" if it mentions the topic but lacks depth, "fail" if it is not about the topic
- reasoning: a brief explanation of your judgment
</system>
<user>
Required topic: {{ required_topic }}
Blog title: {{ blog_title }}
Blog post:
{{ blog_post }}
Judge whether this blog post is faithfully about the required topic.
</user>The eval workflow wires evaluators together and defines how to interpret results.
// tests/evals/workflow.ts
import { evalWorkflow } from '@outputai/evals';
import { evaluateSum } from './evaluators.js';
export default evalWorkflow({
name: 'simple_eval',
evals: [
{
evaluator: evaluateSum,
criticality: 'required',
interpret: { type: 'boolean' }
}
]
});Each entry in the evals array has:
evaluator — the function created by verify()criticality — 'required' (affects pass/fail) or 'informational' (reported but doesn't block)interpret — how to convert the evaluator's return value into a verdict| Type | Evaluator Returns | Mapping |
|---|---|---|
{ type: 'boolean' } | Verdict.equals(), Verdict.gte(), etc. | true = pass, false = fail |
{ type: 'verdict' } | judgeVerdict() or Verdict.pass/partial/fail() | Direct pass-through |
{ type: 'number', pass: 0.7, partial: 0.4 } | judgeScore() | >=pass = pass, >=partial = partial, else fail |
{ type: 'string', pass: ['a', 'b'], partial: ['c'] } | judgeLabel() | Label in pass list = pass, in partial list = partial, else fail |
export default evalWorkflow({
name: 'blog_generator_eval',
evals: [
{
evaluator: lengthOfOutput,
criticality: 'required',
interpret: { type: 'boolean' }
},
{
evaluator: evaluateTopic,
criticality: 'required',
interpret: { type: 'verdict' }
},
{
evaluator: evaluateQuality,
criticality: 'required',
interpret: { type: 'number', pass: 0.7, partial: 0.4 }
},
{
evaluator: evaluateContent,
criticality: 'informational',
interpret: { type: 'boolean' }
},
{
evaluator: evaluateTone,
criticality: 'informational',
interpret: { type: 'string', pass: ['professional', 'informative'], partial: ['casual'] }
}
]
});The eval workflow name must end in _eval and match the pattern {workflow_name}_eval. The CLI resolves this automatically — output workflow test blog_generator looks for blog_generator_eval.
For methodology on designing diverse datasets that cover failure-prone regions, see output-eval-dataset-design.
Datasets are YAML files in tests/datasets/. Each file represents one test case.
name: basic_input
input:
values:
- 1
- 2
- 3
- 4
- 5
last_output:
output:
result: 15
executionTimeMs: 100
date: '2026-02-13T00:00:00.000Z'Ground truth provides expected values for evaluators. You can set global values and per-evaluator overrides:
name: stripe_blog
input:
topic: "Stripe the payment processor"
requirements: "Include a link to https://stripe.com/en-gb/pricing"
last_output:
output:
title: "Stripe: The Modern Payment Processing Platform"
blog_post: |
Stripe has revolutionized online payment processing...
executionTimeMs: 5000
date: '2026-02-16T00:00:00.000Z'
ground_truth:
notes: "Known good case"
evals:
length_of_output:
min_length: 100
evaluate_topic:
required_topic: "Stripe the payment processor"
evaluate_content:
required_content: "https://stripe.com/en-gb/pricing"The ground_truth.evals.<evaluator_name> values are merged with the top-level ground truth and passed to the evaluator via context.ground_truth.
output workflow test <workflow_name>Runs evaluations against all datasets for a workflow.
| Flag | Description |
|---|---|
--cached | Use cached output from dataset files (skip workflow execution) |
--save | Run workflow fresh and save output + eval results back to dataset files |
--dataset <names> | Comma-separated list of dataset names to run (default: all) |
--json | Output machine-readable JSON instead of the rendered report |
Execution flow:
tests/datasets/--cached: executes the workflow for each dataset to get fresh output{workflow_name}_eval workflowoutput workflow dataset list <workflow_name>Lists all datasets for a workflow with their cached status.
| Flag | Description |
|---|---|
--format <type> | Output format: table (default) or text |
--json | Output machine-readable JSON |
output workflow dataset generate <workflow_name> [scenario]Generates a new dataset file by running the workflow.
| Flag | Description |
|---|---|
--input <json> | Workflow input as a JSON string or file path |
--name <name> | Dataset filename (defaults to scenario name) |
--trace <path> | Generate from a local trace file instead of running the workflow |
--download | Download traces from S3 and convert to datasets |
--limit <n> | Max traces to download from S3 (default: 5) |
# Generate dataset from inline JSON input
output workflow dataset generate my_workflow --input '{"key": "value"}' --name my_test
# Generate from a scenario file
output workflow dataset generate my_workflow basic
# Run evals with cached output (fast, no re-execution)
output workflow test my_workflow --cached
# Run evals fresh and save results
output workflow test my_workflow --save
# Run specific datasets only
output workflow test my_workflow --dataset happy_path,edge_case
# List all datasets
output workflow dataset list my_workflow# 1. Start the dev server
npm run output:dev
# 2. Generate datasets from real workflow runs
output workflow dataset generate blog_generator --input '{"topic": "AI"}' --name ai_post
# 3. Edit the dataset YAML to add ground_truth values for your evaluators
# 4. Run evals with --save to cache output and eval results
output workflow test blog_generator --save
# 5. Iterate on evaluators, re-run with cached output (fast)
output workflow test blog_generator --cached
# 6. List all datasets
output workflow dataset list blog_generatorverify, Verdict from @outputai/evals (not @outputai/core)evalWorkflow from @outputai/evals.js extension{workflow_name}_eval patterntests/datasets/tests/evals/name in snake_casecriticality is set to 'required' or 'informational' for each evalinterpret type matches evaluator return type.prompt files are in tests/evals/ alongside evaluatorsz is imported from @outputai/core (not zod)output-dev-evaluator-function — Runtime evaluators using evaluator() from @outputai/coreoutput-dev-scenario-file — Creating scenario JSON files for workflow executionoutput-dev-folder-structure — Understanding project directory layoutoutput-dev-prompt-file — Creating .prompt files for LLM operationsoutput-eval-error-analysis — Identify failure modes before building evaluatorsoutput-eval-judge-prompt — Design effective LLM judge promptsoutput-eval-dataset-design — Generate diverse test datasetsoutput-eval-validate-judge — Validate LLM judges against human labelsoutput-eval-audit — Audit an existing eval suite for trustworthiness© growthxai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing of growthxai/output.
Open the folder on GitHubat commit 52b51ac
Output Dev Eval Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Output Dev Eval Testing this skillgrowthxai/output | 440 | — | ~3.8k | Automated safety check: Notes | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
growthxai/output
Zod schema constraints that Anthropic rejects or silently ignores when sent as structured-output tool definitions via aiSdk.Output.object().
growthxai/output
Implement an Output SDK workflow from a plan document. An agent skill from growthxai/output.
growthxai/output
View, edit, and set encrypted credentials in an Output.ai project.
growthxai/output
Wire encrypted credentials to environment variables using the credential: convention.
growthxai/output
Initialize encrypted credentials for an Output.ai project. An agent skill from growthxai/output.
growthxai/output
Debug Output SDK workflow issues. An agent skill from growthxai/output.
Categories
Create offline evaluation tests for Output SDK workflows using @outputai/evals. Output Dev Eval Testing is an agent skill from growthxai/output. Create offline evaluation tests for Output SDK workflows using @outputai/evals.
Output Dev Eval Testing fits situations like: implementing test evaluators with verify(); creating dataset YAML files; building eval workflows; running workflow tests via CLI.
Run `npx skills add growthxai/output --skill output-dev-eval-testing -a claude-code`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing in growthxai/output) into .claude/skills/output-dev-eval-testing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add growthxai/output --skill output-dev-eval-testing -a codex`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-dev-eval-testing in growthxai/output) into .agents/skills/output-dev-eval-testing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add growthxai/output --skill output-dev-eval-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/output-dev-eval-testing, .gemini/skills/output-dev-eval-testing, .github/skills/output-dev-eval-testing and .opencode/skills/output-dev-eval-testing in your project.
Going by SKILL.md and its folder, Output Dev Eval Testing needs the command-line tools its instructions call (npm). Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.
SKILL.md names 1 domain. In commands or code: stripe.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Output Dev Eval Testing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Output Dev Eval Testing: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
growthxai (a GitHub organization) maintains it in growthxai/output, which has 440 GitHub stars. The repository holds 52 skills in this directory. The repository was last updated on October 7, 2026.
Source: growthxai/output on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.