Perfup
raullenchai/Rapid-MLX
Autonomous performance optimization: research, PoC, benchmark, implement, review, PR
Run and review the Promptfoo-based AFM agentic evaluation suite.
$ npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install scouzi1966/maclocal-api codex-promptfoo-agentic-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/codex-promptfoo-agentic-eval .claude/skills/codex-promptfoo-agentic-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "codex-promptfoo-agentic-eval" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.codex/skills/codex-promptfoo-agentic-eval into .claude/skills/codex-promptfoo-agentic-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codex-promptfoo-agentic-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/scouzi1966/maclocal-api/tree/main/.codex/skills/codex-promptfoo-agentic-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install scouzi1966/maclocal-api codex-promptfoo-agentic-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.codex/skills/codex-promptfoo-agentic-eval .agents/skills/codex-promptfoo-agentic-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "codex-promptfoo-agentic-eval" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.codex/skills/codex-promptfoo-agentic-eval into .agents/skills/codex-promptfoo-agentic-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codex-promptfoo-agentic-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install scouzi1966/maclocal-api codex-promptfoo-agentic-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.codex/skills/codex-promptfoo-agentic-eval .cursor/skills/codex-promptfoo-agentic-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "codex-promptfoo-agentic-eval" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.codex/skills/codex-promptfoo-agentic-eval into .cursor/skills/codex-promptfoo-agentic-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codex-promptfoo-agentic-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/scouzi1966/maclocal-api.git --path .codex/skills/codex-promptfoo-agentic-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install scouzi1966/maclocal-api codex-promptfoo-agentic-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.codex/skills/codex-promptfoo-agentic-eval .gemini/skills/codex-promptfoo-agentic-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "codex-promptfoo-agentic-eval" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.codex/skills/codex-promptfoo-agentic-eval into .gemini/skills/codex-promptfoo-agentic-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codex-promptfoo-agentic-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install scouzi1966/maclocal-api codex-promptfoo-agentic-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .github/skills && cp -r skills-src/.codex/skills/codex-promptfoo-agentic-eval .github/skills/codex-promptfoo-agentic-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "codex-promptfoo-agentic-eval" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.codex/skills/codex-promptfoo-agentic-eval into .github/skills/codex-promptfoo-agentic-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codex-promptfoo-agentic-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install scouzi1966/maclocal-api codex-promptfoo-agentic-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.codex/skills/codex-promptfoo-agentic-eval .opencode/skills/codex-promptfoo-agentic-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "codex-promptfoo-agentic-eval" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.codex/skills/codex-promptfoo-agentic-eval into .opencode/skills/codex-promptfoo-agentic-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codex-promptfoo-agentic-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
codex-promptfoo-agentic-evalRun and review the Promptfoo-based AFM agentic evaluation suite.
Codex Promptfoo Agentic Eval is an agent skill from scouzi1966/maclocal-api. Run and review the Promptfoo-based AFM agentic evaluation suite. Use when the user wants structured-output, tool-calling, grammar, guided-json, streaming, concurrency, or agentic QA coverage for AFM, and especially when they want help choosing harness options or interpreting failures.
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Structured output and tool calling. It works with macOS. The repository describes itself as: 'afm' command cli: macOS server and single prompt mode that exposes Apple's Foundation and MLX Models and other APIs running on your Mac through a single aggregated…. The licence is MIT.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 138ca5d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
nodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Codex Promptfoo Agentic Eval loads about 1.8k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 587 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from scouzi1966/maclocal-api at commit 138ca5d, republished under its MIT licence (© scouzi1966). 587 words, ~1,845 tokens.
.claude/skills/codex-promptfoo-agentic-eval/SKILL.md (or your agent's skills folder).Use this skill when the user wants to run, expand, or interpret the Promptfoo agentic suite for AFM.
This skill is for two linked goals:
Always distinguish:
afm_bugmodel_qualityharness_bugAlways report provenance for the suite you run:
afm_internalprimary_sourcepublic_benchmark_inspiredsyntheticNever present a benchmark-inspired or synthetic suite as if it were a public benchmark import.
Before running the suite, ask the user the minimum needed questions:
structuredstructured-stresstoolcalltoolcall-qualityagenticframeworksopencodealldefault, adaptive-xml, adaptive-xml-grammarIf the user does not specify, assume:
allbothRun from:
cd /Volumes/edata/codex/dev/git/maclocal-api/NEXT/maclocal-apiRead only what is needed:
Scripts/feature-promptfoo-agentic/README.mdScripts/feature-promptfoo-agentic/run-promptfoo-agentic.shScripts/feature-promptfoo-agentic/providers/afm_provider.mjsScripts/feature-promptfoo-agentic/matrix/functional-matrix.yamlScripts/feature-promptfoo-agentic/matrix/failure-classification.yamldocs/roadmap/promptfoo-agentic-matrix.mdRelevant suite configs and datasets:
Scripts/feature-promptfoo-agentic/promptfooconfig.structured.yamlScripts/feature-promptfoo-agentic/promptfooconfig.structured-stress.yamlScripts/feature-promptfoo-agentic/promptfooconfig.toolcall.yamlScripts/feature-promptfoo-agentic/promptfooconfig.toolcall-quality.yamlScripts/feature-promptfoo-agentic/promptfooconfig.agentic.yamlScripts/feature-promptfoo-agentic/promptfooconfig.agentic-frameworks.yamlScripts/feature-promptfoo-agentic/promptfooconfig.opencode.yamlScripts/feature-promptfoo-agentic/datasets/agentic/opencode-primary-tools.yamlIf reviewing failures, inspect:
test-reports/promptfoo-agentic/*.jsontest-reports/promptfoo-agentic/*.classified.jsontest-reports/promptfoo-agentic/*.classified.summary.mdtest-reports/promptfoo-agentic/server-*.logUse the wrapper unless the user explicitly wants a narrower manual run:
MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache \
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.shAllowed narrowed runs:
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh structured
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh structured-stress
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh toolcall
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh toolcall-quality
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh agentic
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh frameworks
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh opencode
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh default
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh adaptive-xml
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh adaptive-xml-grammarKnown current suite sizes:
structured: 6 cases (afm_internal)structured-stress: 4 cases (public_benchmark_inspired)toolcall: 7 cases (afm_internal)toolcall-quality: 6 cases (public_benchmark_inspired)agentic: 4 cases (synthetic representative)frameworks: 8 cases (mixed / currently assumption-heavy)opencode: 37 cases (primary_source)If a requested run is under 20 cases, explicitly warn the user that it is a small sample, not broad coverage.
Check:
If using the automated judge path:
AFM_JUDGE_MODEL="$MODEL_ID" \
AFM_JUDGE_BASE_URL=http://127.0.0.1:9999/v1 \
node Scripts/feature-promptfoo-agentic/judges/classify-failures.mjs <report.json>If working interactively in Codex/Claude-style CLI, classify manually using the rubric in:
docs/roadmap/promptfoo-agentic-matrix.mdScripts/feature-promptfoo-agentic/matrix/failure-classification.yamlafm_bugUse when AFM violates server/runtime/protocol invariants:
tool_calls envelopetool_choice semanticsmodel_qualityUse when AFM output is valid but the model behavior is weak:
harness_bugUse when the test machinery is wrong:
When reporting results, give:
afm_internalprimary_sourcepublic_benchmark_inspiredsyntheticafm_bugmodel_qualityharness_bugnot_yet_classified count, if anyPrefer concise summaries, but include concrete failing cases when they matter.
When the user asks to extend the suite, prioritize:
Prefer primary sources over secondary descriptions. If a suite is built from secondary material or assumptions, say so explicitly and do not overstate its authority.
Do not explode the matrix blindly. Use the layered matrix in
docs/roadmap/promptfoo-agentic-matrix.md.
© scouzi1966, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .codex/skills/codex-promptfoo-agentic-eval of scouzi1966/maclocal-api.
Open the folder on GitHubat commit 138ca5d
Codex Promptfoo Agentic Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Codex Promptfoo Agentic Eval this skillscouzi1966/maclocal-api | 345 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Perfupraullenchai/Rapid-MLX | 3.9k | — | ~1.6k | Automated safety check: Notes | Custom licence | |
| Foundation Models App Builderrryam/FoundationModelsKit | 162 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Planning With Filesjarrodwatts/claude-code-config | 1.1k | 5 repos | ~967 | Automated safety check: Pass | None | |
| Tool Use Data Synthesissunny-glow/Auto-BenchMax | 1.3k | — | ~3.3k | Automated safety check: Pass | None | |
| Agent Harness ConstructionKartikLabhshetwar/mind-mentor | 147 | 7 repos | ~500 | Automated safety check: Pass | Apache-2.0 |
raullenchai/Rapid-MLX
Autonomous performance optimization: research, PoC, benchmark, implement, review, PR
rryam/FoundationModelsKit
Build or modify Apple Foundation Models features in Swift, SwiftUI, iOS, and macOS apps.
jarrodwatts/claude-code-config
Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.
sunny-glow/Auto-BenchMax
Synthesize training data for ANY tool-use / agentic benchmark, in ANY repo.
KartikLabhshetwar/mind-mentor
Design and optimize AI agent action spaces, tool definitions, and observation formatting for higher completion rates.
wshobson/agents
Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.
scouzi1966/maclocal-api
Maintain and extend AFM (maclocal-api), a Swift OpenAI-compatible local LLM server and CLI for Apple Foundation Models, MLX models, API gateway proxying, and Vision OCR.
scouzi1966/maclocal-api
Build AFM from scratch — submodules, patches, webui, and Swift build.
scouzi1966/maclocal-api
Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…
scouzi1966/maclocal-api
A skill your agent uses when user wants to build a PyPI wheel from an existing compiled afm binary and publish to PyPI.
scouzi1966/maclocal-api
Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.
scouzi1966/maclocal-api
A skill your agent uses when testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors, classifies them as afm translation bugs vs model generation errors, and…
Works with
Categories
Run and review the Promptfoo-based AFM agentic evaluation suite. Codex Promptfoo Agentic Eval is an agent skill from scouzi1966/maclocal-api. Run and review the Promptfoo-based AFM agentic evaluation suite.
Codex Promptfoo Agentic Eval fits situations like: the user wants structured-output; agentic QA coverage for AFM; especially when they want help choosing harness options; interpreting failures.
Run `npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a claude-code`. Or copy the skill folder (.codex/skills/codex-promptfoo-agentic-eval in scouzi1966/maclocal-api) into .claude/skills/codex-promptfoo-agentic-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a codex`. Or copy the skill folder (.codex/skills/codex-promptfoo-agentic-eval in scouzi1966/maclocal-api) into .agents/skills/codex-promptfoo-agentic-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/codex-promptfoo-agentic-eval, .gemini/skills/codex-promptfoo-agentic-eval, .github/skills/codex-promptfoo-agentic-eval and .opencode/skills/codex-promptfoo-agentic-eval in your project.
Going by SKILL.md and its folder, Codex Promptfoo Agentic Eval needs the command-line tools its instructions call (node).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Codex Promptfoo Agentic Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Codex Promptfoo Agentic Eval: Perfup (raullenchai/Rapid-MLX, 3.9k stars), Foundation Models App Builder (rryam/FoundationModelsKit, 162 stars), Planning With Files (jarrodwatts/claude-code-config, 1.1k stars) and Tool Use Data Synthesis (sunny-glow/Auto-BenchMax, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
scouzi1966 (a GitHub user) maintains it in scouzi1966/maclocal-api, which has 345 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 5, 2026.
Source: scouzi1966/maclocal-api on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.