Agent skill

Codex Promptfoo Agentic Eval

by scouzi1966 in scouzi1966/maclocal-api

Run and review the Promptfoo-based AFM agentic evaluation suite.

MITAuto-check passedAI & LLM Engineering

Install Codex Promptfoo Agentic Eval

skills CLI
$ npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scouzi1966/maclocal-api codex-promptfoo-agentic-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/codex-promptfoo-agentic-eval .claude/skills/codex-promptfoo-agentic-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
codex-promptfoo-agentic-eval
GitHub stars
345
Token cost
~1.8k tokens
SKILL.md length
587 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Run and review the Promptfoo-based AFM agentic evaluation suite.

  • Works in 3 steps: Run the harness → Review results → Classify failures
  • The user wants structured-output
  • SKILL.md covers First questions to ask, Working directory, Main assets and Execution workflow, plus 3 more sections
  • Calls node

What it does

Codex Promptfoo Agentic Eval is an agent skill from scouzi1966/maclocal-api. Run and review the Promptfoo-based AFM agentic evaluation suite. Use when the user wants structured-output, tool-calling, grammar, guided-json, streaming, concurrency, or agentic QA coverage for AFM, and especially when they want help choosing harness options or interpreting failures.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Structured output and tool calling. It works with macOS. The repository describes itself as: 'afm' command cli: macOS server and single prompt mode that exposes Apple's Foundation and MLX Models and other APIs running on your Mac through a single aggregated…. The licence is MIT.

When your agent uses it

  • The user wants structured-output
  • Agentic QA coverage for AFM
  • Especially when they want help choosing harness options
  • Interpreting failures

Example prompts

  • “/codex-promptfoo-agentic-eval”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Run the harness
  2. Review results
  3. Classify failures

What it can do on your machine

Read from SKILL.md and the folder at commit 138ca5d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Codex Promptfoo Agentic Eval loads about 1.8k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 587 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scouzi1966/maclocal-api at commit 138ca5d, republished under its MIT licence (© scouzi1966). 587 words, ~1,845 tokens.

Download SKILL.mdSave it as .claude/skills/codex-promptfoo-agentic-eval/SKILL.md (or your agent's skills folder).
name
codex-promptfoo-agentic-eval
description
Run and review the Promptfoo-based AFM agentic evaluation suite. Use when the user wants structured-output, tool-calling, grammar, guided-json, streaming, concurrency, or agentic QA coverage for AFM, and especially when they want help choosing harness options or interpreting failures.

Promptfoo Agentic Eval

Use this skill when the user wants to run, expand, or interpret the Promptfoo agentic suite for AFM.

This skill is for two linked goals:

  • AFM functional validation
  • model-quality evaluation for agentic use

Always distinguish:

  • afm_bug
  • model_quality
  • harness_bug

Always report provenance for the suite you run:

  • afm_internal
  • primary_source
  • public_benchmark_inspired
  • synthetic

Never present a benchmark-inspired or synthetic suite as if it were a public benchmark import.

First questions to ask

Before running the suite, ask the user the minimum needed questions:

  1. Which model should be tested?
  2. Which scope should be run?
    • structured
    • structured-stress
    • toolcall
    • toolcall-quality
    • agentic
    • frameworks
    • opencode
    • all
    • one profile only: default, adaptive-xml, adaptive-xml-grammar
  3. Is the goal:
    • AFM functional QA
    • model quality
    • both
  4. Should the run stay serial/safe, or include concurrency cases?
  5. Should you only review existing reports, or also execute the harness?
  6. Should the run prefer:
    • primary-source-only cases
    • public-benchmark-inspired cases
    • synthetic representative cases
    • mixed

If the user does not specify, assume:

  • model: the repo's current primary MLX model under test
  • scope: all
  • goal: both
  • run mode: serial/safe first
  • action: execute and then review
  • provenance preference: mixed, but explicitly labeled

Working directory

Run from:

bash
cd /Volumes/edata/codex/dev/git/maclocal-api/NEXT/maclocal-api

Main assets

Read only what is needed:

  • Scripts/feature-promptfoo-agentic/README.md
  • Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh
  • Scripts/feature-promptfoo-agentic/providers/afm_provider.mjs
  • Scripts/feature-promptfoo-agentic/matrix/functional-matrix.yaml
  • Scripts/feature-promptfoo-agentic/matrix/failure-classification.yaml
  • docs/roadmap/promptfoo-agentic-matrix.md

Relevant suite configs and datasets:

  • Scripts/feature-promptfoo-agentic/promptfooconfig.structured.yaml
  • Scripts/feature-promptfoo-agentic/promptfooconfig.structured-stress.yaml
  • Scripts/feature-promptfoo-agentic/promptfooconfig.toolcall.yaml
  • Scripts/feature-promptfoo-agentic/promptfooconfig.toolcall-quality.yaml
  • Scripts/feature-promptfoo-agentic/promptfooconfig.agentic.yaml
  • Scripts/feature-promptfoo-agentic/promptfooconfig.agentic-frameworks.yaml
  • Scripts/feature-promptfoo-agentic/promptfooconfig.opencode.yaml
  • Scripts/feature-promptfoo-agentic/datasets/agentic/opencode-primary-tools.yaml

If reviewing failures, inspect:

  • test-reports/promptfoo-agentic/*.json
  • test-reports/promptfoo-agentic/*.classified.json
  • test-reports/promptfoo-agentic/*.classified.summary.md
  • test-reports/promptfoo-agentic/server-*.log

Execution workflow

1. Run the harness

Use the wrapper unless the user explicitly wants a narrower manual run:

bash
MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache \
  Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh

Allowed narrowed runs:

bash
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh structured
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh structured-stress
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh toolcall
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh toolcall-quality
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh agentic
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh frameworks
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh opencode
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh default
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh adaptive-xml
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh adaptive-xml-grammar

Known current suite sizes:

  • structured: 6 cases (afm_internal)
  • structured-stress: 4 cases (public_benchmark_inspired)
  • toolcall: 7 cases (afm_internal)
  • toolcall-quality: 6 cases (public_benchmark_inspired)
  • agentic: 4 cases (synthetic representative)
  • frameworks: 8 cases (mixed / currently assumption-heavy)
  • opencode: 37 cases (primary_source)

If a requested run is under 20 cases, explicitly warn the user that it is a small sample, not broad coverage.

2. Review results

Check:

  • pass/fail counts
  • whether failures are real or harness-related
  • differences across parser profiles
  • structured-output vs tool-calling behavior
  • provenance of the suite and whether that limits how strong conclusions can be
3. Classify failures

If using the automated judge path:

bash
AFM_JUDGE_MODEL="$MODEL_ID" \
AFM_JUDGE_BASE_URL=http://127.0.0.1:9999/v1 \
node Scripts/feature-promptfoo-agentic/judges/classify-failures.mjs <report.json>

If working interactively in Codex/Claude-style CLI, classify manually using the rubric in:

  • docs/roadmap/promptfoo-agentic-matrix.md
  • Scripts/feature-promptfoo-agentic/matrix/failure-classification.yaml
Show full SKILL.md (227 more words)Show less

Classification rubric

afm_bug

Use when AFM violates server/runtime/protocol invariants:

  • malformed JSON or SSE
  • broken tool_calls envelope
  • wrong tool_choice semantics
  • grammar-constrained output violates grammar/schema
  • stream/non-stream deterministic mismatch
  • parser corrupts an otherwise valid call
  • timeout, truncation, duplicate emission, crash
model_quality

Use when AFM output is valid but the model behavior is weak:

  • wrong tool
  • missing tool
  • unnecessary tool
  • wrong arguments
  • poor multi-turn or refusal behavior
harness_bug

Use when the test machinery is wrong:

  • assertion false negative
  • provider normalization issue
  • Promptfoo config mismatch
  • classification/judge pipeline issue

Reporting format

When reporting results, give:

  1. overall run status
  2. total tests executed
  3. pass/fail counts per suite/profile
  4. suite provenance summary
    • how many cases came from afm_internal
    • primary_source
    • public_benchmark_inspired
    • synthetic
  5. failure classification summary:
    • afm_bug
    • model_quality
    • harness_bug
  6. remaining not_yet_classified count, if any
  7. top next actions

Prefer concise summaries, but include concrete failing cases when they matter.

Expansion guidance

When the user asks to extend the suite, prioritize:

  1. stronger custom assertions
  2. streaming and grammar-specific cases
  3. primary-source-derived framework suites
  4. public benchmark sampling:
    • BFCL
    • When2Call
    • StructEval
    • tau-bench-style multi-turn cases
  5. real AFM use cases:
    • coding agents
    • OpenClaw/Hermes-style tool orchestration
    • structured output workflows

Prefer primary sources over secondary descriptions. If a suite is built from secondary material or assumptions, say so explicitly and do not overstate its authority.

Do not explode the matrix blindly. Use the layered matrix in docs/roadmap/promptfoo-agentic-matrix.md.

© scouzi1966, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .codex/skills/codex-promptfoo-agentic-eval of scouzi1966/maclocal-api.

Open the folder on GitHubat commit 138ca5d

Compare with similar skills

Codex Promptfoo Agentic Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Codex Promptfoo Agentic Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Codex Promptfoo Agentic Eval this skillscouzi1966/maclocal-api345—~1.8kAutomated safety check: PassMIT
Perfupraullenchai/Rapid-MLX3.9k—~1.6kAutomated safety check: NotesCustom licence
Foundation Models App Builderrryam/FoundationModelsKit162—~1.1kAutomated safety check: PassMIT
Planning With Filesjarrodwatts/claude-code-config1.1k5 repos~967Automated safety check: PassNone
Tool Use Data Synthesissunny-glow/Auto-BenchMax1.3k—~3.3kAutomated safety check: PassNone
Agent Harness ConstructionKartikLabhshetwar/mind-mentor1477 repos~500Automated safety check: PassApache-2.0

Similar skills

  • Perfup

    raullenchai/Rapid-MLX

    Autonomous performance optimization: research, PoC, benchmark, implement, review, PR

    3.9k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Foundation Models App Builder

    rryam/FoundationModelsKit

    Build or modify Apple Foundation Models features in Swift, SwiftUI, iOS, and macOS apps.

    162 GitHub stars~1.1k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Planning With Files

    jarrodwatts/claude-code-config

    Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.

    1.1k GitHub starsUsed in 5 repos~967 tokens
    AI & LLM EngineeringAuto-check passed
  • Tool Use Data Synthesis

    sunny-glow/Auto-BenchMax

    Synthesize training data for ANY tool-use / agentic benchmark, in ANY repo.

    1.3k GitHub stars~3.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agent Harness Construction

    KartikLabhshetwar/mind-mentor

    Design and optimize AI agent action spaces, tool definitions, and observation formatting for higher completion rates.

    147 GitHub starsUsed in 7 repos~500 tokens
    AI & LLM EngineeringAuto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from scouzi1966/maclocal-api

All 12 skills in this repo
  • Afm

    scouzi1966/maclocal-api

    Maintain and extend AFM (maclocal-api), a Swift OpenAI-compatible local LLM server and CLI for Apple Foundation Models, MLX models, API gateway proxying, and Vision OCR.

    345 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • Build Afm

    scouzi1966/maclocal-api

    Build AFM from scratch — submodules, patches, webui, and Swift build.

    345 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check: notes
  • Test Afm Binary

    scouzi1966/maclocal-api

    Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…

    345 GitHub stars~3.8k tokensUpdated 2 days ago
    Auto-check passed
  • Afm Release Wheel

    scouzi1966/maclocal-api

    A skill your agent uses when user wants to build a PyPI wheel from an existing compiled afm binary and publish to PyPI.

    345 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check: warnings
  • Test Macafm

    scouzi1966/maclocal-api

    Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.

    345 GitHub stars~7k tokensUpdated 2 days ago
    Auto-check passed
  • Test Opencode Tooling

    scouzi1966/maclocal-api

    A skill your agent uses when testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors, classifies them as afm translation bugs vs model generation errors, and…

    345 GitHub stars~4.2k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Codex Promptfoo Agentic Eval

What does Codex Promptfoo Agentic Eval do?

Run and review the Promptfoo-based AFM agentic evaluation suite. Codex Promptfoo Agentic Eval is an agent skill from scouzi1966/maclocal-api. Run and review the Promptfoo-based AFM agentic evaluation suite.

When should I use Codex Promptfoo Agentic Eval?

Codex Promptfoo Agentic Eval fits situations like: the user wants structured-output; agentic QA coverage for AFM; especially when they want help choosing harness options; interpreting failures.

How do I install Codex Promptfoo Agentic Eval in Claude Code?

Run `npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a claude-code`. Or copy the skill folder (.codex/skills/codex-promptfoo-agentic-eval in scouzi1966/maclocal-api) into .claude/skills/codex-promptfoo-agentic-eval in your project. Claude Code loads it when a task matches its description.

How do I install Codex Promptfoo Agentic Eval in Codex?

Run `npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a codex`. Or copy the skill folder (.codex/skills/codex-promptfoo-agentic-eval in scouzi1966/maclocal-api) into .agents/skills/codex-promptfoo-agentic-eval in your project. Codex loads it when a task matches its description.

Can I use Codex Promptfoo Agentic Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/codex-promptfoo-agentic-eval, .gemini/skills/codex-promptfoo-agentic-eval, .github/skills/codex-promptfoo-agentic-eval and .opencode/skills/codex-promptfoo-agentic-eval in your project.

What does Codex Promptfoo Agentic Eval need to run?

Going by SKILL.md and its folder, Codex Promptfoo Agentic Eval needs the command-line tools its instructions call (node).

Does Codex Promptfoo Agentic Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Codex Promptfoo Agentic Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Codex Promptfoo Agentic Eval use?

Codex Promptfoo Agentic Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Codex Promptfoo Agentic Eval use?

About 1.8k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Codex Promptfoo Agentic Eval?

Skills that share tags, products or a category with Codex Promptfoo Agentic Eval: Perfup (raullenchai/Rapid-MLX, 3.9k stars), Foundation Models App Builder (rryam/FoundationModelsKit, 162 stars), Planning With Files (jarrodwatts/claude-code-config, 1.1k stars) and Tool Use Data Synthesis (sunny-glow/Auto-BenchMax, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Codex Promptfoo Agentic Eval?

scouzi1966 (a GitHub user) maintains it in scouzi1966/maclocal-api, which has 345 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 5, 2026.

Source: scouzi1966/maclocal-api on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.