Agent Eval Cases
agentailor/fullstack-langgraph-nextjs-agent
Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.
Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…
$ npx skills add scouzi1966/maclocal-api --skill test-afm-binary -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install scouzi1966/maclocal-api test-afm-binary --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/test-afm-binary .claude/skills/test-afm-binary && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "test-afm-binary" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-afm-binary into .claude/skills/test-afm-binary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-afm-binary", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-afm-binaryType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add scouzi1966/maclocal-api --skill test-afm-binary -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install scouzi1966/maclocal-api test-afm-binary --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/test-afm-binary .agents/skills/test-afm-binary && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "test-afm-binary" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-afm-binary into .agents/skills/test-afm-binary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-afm-binary", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scouzi1966/maclocal-api --skill test-afm-binary -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install scouzi1966/maclocal-api test-afm-binary --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/test-afm-binary .cursor/skills/test-afm-binary && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "test-afm-binary" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-afm-binary into .cursor/skills/test-afm-binary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-afm-binary", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/scouzi1966/maclocal-api.git --path .claude/skills/test-afm-binary--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add scouzi1966/maclocal-api --skill test-afm-binary -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install scouzi1966/maclocal-api test-afm-binary --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/test-afm-binary .gemini/skills/test-afm-binary && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "test-afm-binary" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-afm-binary into .gemini/skills/test-afm-binary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-afm-binary", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install scouzi1966/maclocal-api test-afm-binaryInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add scouzi1966/maclocal-api --skill test-afm-binary -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/test-afm-binary .github/skills/test-afm-binary && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "test-afm-binary" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-afm-binary into .github/skills/test-afm-binary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-afm-binary", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scouzi1966/maclocal-api --skill test-afm-binary -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install scouzi1966/maclocal-api test-afm-binary --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/test-afm-binary .opencode/skills/test-afm-binary && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "test-afm-binary" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-afm-binary into .opencode/skills/test-afm-binary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-afm-binary", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
test-afm-binaryTest a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…
Test Afm Binary is an agent skill from scouzi1966/maclocal-api. Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU profiling. Use when user wants to validate a binary post-build, after code changes, or before release.
Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation and Unit testing. It works with OpenAI and macOS. The repository describes itself as: 'afm' command cli: macOS server and single prompt mode that exposes Apple's Foundation and MLX Models and other APIs running on your Mac through a single aggregated…. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 138ca5d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3brewswiftFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Test Afm Binary loads about 3.8k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 1,190 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from scouzi1966/maclocal-api at commit 138ca5d, republished under its MIT licence (© scouzi1966). 1,190 words, ~3,753 tokens.
.claude/skills/test-afm-binary/SKILL.md (or your agent's skills folder).Test any pre-built afm binary with a menu of test suites. Validates the binary won't crash when relocated (pip/Homebrew install), then runs the selected tests.
/test-afm-binary — interactive: asks for binary path, model, and test selection/test-afm-binary /path/to/afm — test the binary at the given path/test-afm-binary .build/arm64-apple-macosx/release/afm — test the current buildAsk for the binary path if not provided as an argument. Default: .build/arm64-apple-macosx/release/afm.
BIN="${1:-.build/arm64-apple-macosx/release/afm}"
[ -x "$BIN" ] || BIN=".build/release/afm"
BIN_ABS="$(cd "$(dirname "$BIN")" && pwd)/$(basename "$BIN")"
echo "Binary: $BIN_ABS"If the binary doesn't exist or isn't executable, STOP and tell the user.
These checks run before any test suite. They catch fatal distribution bugs that would crash every pip/Homebrew user. If any check fails, STOP — do not proceed to testing.
REPORTED=$($BIN_ABS --version 2>&1)
echo "Version: $REPORTED"If the version shows only a base version without a SHA suffix (e.g., v0.9.8 instead of v0.9.8-62395ab), warn the user: this likely means the binary was built with an incremental swift build instead of ./Scripts/build-from-scratch.sh. The SHA injection only happens in the build script. This is a warning, not a blocker — the binary may still be valid for testing.
BIN_DIR="$(dirname "$BIN_ABS")"
# Check for metallib in either location (SPM bundle or loose file)
if [ -f "$BIN_DIR/AFMKit_AFMKitMLX.bundle/default.metallib" ]; then
echo "PASS: Metallib in SPM bundle ($(du -h "$BIN_DIR/AFMKit_AFMKitMLX.bundle/default.metallib" | cut -f1))"
elif [ -f "$BIN_DIR/default.metallib" ]; then
echo "PASS: Loose metallib ($(du -h "$BIN_DIR/default.metallib" | cut -f1))"
else
echo "FAIL: No metallib found next to binary"
echo "The binary will crash on first inference without default.metallib"
fiTMPDIR=$(mktemp -d)
cp "$BIN_ABS" "$TMPDIR/"
# Copy metallib as loose file (pip wheel layout)
if [ -f "$BIN_DIR/AFMKit_AFMKitMLX.bundle/default.metallib" ]; then
cp "$BIN_DIR/AFMKit_AFMKitMLX.bundle/default.metallib" "$TMPDIR/"
elif [ -f "$BIN_DIR/default.metallib" ]; then
cp "$BIN_DIR/default.metallib" "$TMPDIR/"
fi
MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache \
"$TMPDIR/afm" mlx -m mlx-community/SmolLM3-3B-4bit -s "hello" --max-tokens 3 2>&1 | head -3
EXIT_CODE=${PIPESTATUS[0]}
rm -rf "$TMPDIR"
if [ "$EXIT_CODE" -ne 0 ]; then
echo "FATAL: Relocated binary crashed (exit $EXIT_CODE)"
echo "Bundle.module fatalError is still reachable — pip/Homebrew install will crash"
echo "STOP. Fix MLXMetalLibrary.swift — it must NOT call Bundle.module"
else
echo "PASS: Relocated binary runs without crash"
fiIf this fails, STOP IMMEDIATELY. Do not run any tests. The binary is broken for distribution.
HITS=$(grep -r 'Bundle\.module' Sources/ --include='*.swift' | grep -v '^\s*//' | grep -v '// ' | wc -l | tr -d ' ')
if [ "$HITS" -gt 0 ]; then
echo "FAIL: Found $HITS Bundle.module call(s) in source"
grep -rn 'Bundle\.module' Sources/ --include='*.swift' | grep -v '//'
echo "This WILL crash when installed via pip or Homebrew"
else
echo "PASS: No Bundle.module calls in source"
fimacOS 26 SIGABRTs any process that requests privacy-sensitive APIs (Speech Recognition, microphone, camera, etc.) without a matching *UsageDescription key in the binary's embedded Info.plist. PR #107's Apple Speech feature triggers this on every afm speech / POST /v1/audio/transcriptions / chat input_audio call.
# Verify __TEXT,__info_plist section exists
if otool -l "$BIN_ABS" | grep -q '__info_plist'; then
echo "PASS: __info_plist section present"
else
echo "FAIL: Missing __TEXT,__info_plist section"
echo "Check Package.swift linker flags and Sources/AFMCLI/Info.plist"
fi
# Verify NSSpeechRecognitionUsageDescription key is in the embedded plist
if strings "$BIN_ABS" | grep -q 'NSSpeechRecognitionUsageDescription'; then
echo "PASS: NSSpeechRecognitionUsageDescription embedded"
else
echo "FAIL: NSSpeechRecognitionUsageDescription missing"
echo "afm speech / /v1/audio/transcriptions will SIGABRT on macOS 26"
fiNote on testing Speech from an unattended context: If this skill is running inside Claude Code / an editor terminal / any parent process that does NOT have NSSpeechRecognitionUsageDescription, macOS 26 attributes the TCC subject to the parent and the child crashes even with a correct embedded plist. This is a test-environment artifact, not a binary bug. To verify Speech end-to-end, run afm speech -f <file.wav> from a fresh Terminal.app window (stock /System/Applications/Utilities/Terminal.app).
| Check | What it catches | Result |
|---|---|---|
| A: Version | Incremental build (no SHA) | PASS/WARN/FAIL |
| B: Metallib | Missing Metal shaders → crash on inference | PASS/FAIL |
| C: Relocated binary | Bundle.module fatalError → crash on pip install | PASS/FAIL |
| D: No Bundle.module | Source code regression guard | PASS/FAIL |
| E: Info.plist embedded | macOS 26 SIGABRT on Speech Recognition without UsageDescription | PASS/FAIL |
If B, C, D, or E fail, STOP. Do not proceed.
Show available models and let the user pick:
MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache ./Scripts/list-models.shUse AskUserQuestion with the model list. Default recommendation: mlx-community/Qwen3.5-35B-A3B-4bit (19 GB, MoE, best coverage).
For quick smoke tests, suggest mlx-community/SmolLM3-3B-4bit (1.6 GB, fast).
Use AskUserQuestion with multiSelect: true. Present these options:
| Option | Script | Server? | Port | Runtime | What it tests |
|---|---|---|---|---|---|
| All | (runs everything below) | — | — | ~3-4 hours | Complete validation |
| Unit tests | Scripts/swiftpm-reliable.sh test | No | — | ~5s | Swift unit tests (XML parsing, batch scheduler, KV cache, etc.) |
| Assertions (smoke) | test-assertions.sh --tier smoke | Yes | 9998 | ~2 min | Server reachable, basic completion, stop, logprobs, think, tools, errors |
| Assertions (standard) | test-assertions.sh --tier standard | Yes | 9998 | ~5 min | + streaming, cache, concurrent, kwargs, XML tools, adaptive XML, grammar, batch |
| Assertions (full) | test-assertions.sh --tier full | Yes | 9998 | ~15 min | + performance (TTFT, tok/s, long context 2K/4K tokens) |
| Assertions + grammar + forced parser | test-assertions-multi.sh | Managed | 9998 | ~30 min | Full tier × 2 (auto-detect + forced qwen3_xml) with grammar constraints |
| Comprehensive smart analysis | mlx-model-test.sh --smart 1:claude | Managed | 9877 | ~45-90 min | 91 test variants across samplers, stop, JSON, tools, code, math with AI judge |
| Promptfoo agentic evals | run-promptfoo-agentic.sh all | Managed | 9999 | ~60-120 min | 137 tests × 8 server profiles: structured, toolcall, grammar, agentic, frameworks |
| Batch correctness | validate_responses.py | Yes | 9999 | ~10-15 min | Known-answer correctness at B={1,2,4,8} |
| Batch mixed workload | validate_mixed_workload.py | Yes | 9999 | ~15-25 min | Short+long decode mix with GPU metrics |
| Batch multiturn prefix | validate_multiturn_prefix.py | Yes | 9999 | ~15-25 min | Multi-turn prefix cache under concurrency |
| OpenAI compat evals | test-openai-compat-evals.py | Managed | 9999 | ~5-10 min | OpenAI Python SDK compatibility (stream, logprobs, usage) |
| Guided JSON evals | test-guided-json-evals.py | Managed | 9999 | ~10-15 min | response_format: json_schema with real-world fixtures |
| GPU profile | gpu-profile-report.py | No (CLI) | — | ~30-60s | DRAM bandwidth, GPU power, shader kernel names, HTML report |
For each selected test, set the correct environment and invoke. The binary path must be passed to every script.
Environment (always set):
export MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cacheParallelism rules:
Per-test invocation:
| Test | Command |
|---|---|
| Unit tests | Scripts/swiftpm-reliable.sh test |
| Assertions (any tier) | Start server: MACAFM_MLX_MODEL_CACHE=... $BIN_ABS mlx -m MODEL --port 9998 --tool-call-parser afm_adaptive_xml --enable-prefix-caching --enable-grammar-constraints & then ./Scripts/test-assertions.sh --tier TIER --model MODEL --port 9998 --bin "$BIN_ABS" --grammar-constraints |
| Assertions + grammar + forced | ./Scripts/test-assertions-multi.sh --models "MODEL" --tier full --also-forced-parser qwen3_xml --grammar-constraints with AFM_BINARY="$BIN_ABS" |
| Smart analysis | AFM_BIN="$BIN_ABS" ./Scripts/mlx-model-test.sh --model MODEL --prompts Scripts/test-llm-comprehensive.txt --smart 1:claude |
| Promptfoo | AFM_MODEL=MODEL AFM_BINARY="$BIN_ABS" ./Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh all |
| Batch correctness | Start server: $BIN_ABS mlx -m MODEL --port 9999 --concurrent 8 & then python3 Scripts/feature-mlx-concurrent-batch/validate_responses.py |
| Batch mixed | Same server, then python3 Scripts/feature-mlx-concurrent-batch/validate_mixed_workload.py |
| Batch multiturn | Same server, then python3 Scripts/feature-mlx-concurrent-batch/validate_multiturn_prefix.py |
| OpenAI compat | python3 Scripts/feature-codex-optimize-api/test-openai-compat-evals.py --start-server --model MODEL with AFM_BINARY="$BIN_ABS" |
| Guided JSON | python3 Scripts/feature-codex-optimize-api/test-guided-json-evals.py --start-server --model MODEL with AFM_BINARY="$BIN_ABS" |
| GPU profile | python3 Scripts/gpu-profile-report.py MODEL with AFM_BIN="$BIN_ABS" |
After each test completes, present its results immediately. Don't wait for all tests to finish before showing anything.
After all selected tests complete, present a summary table:
| Suite | Pass | Total | Rate | Notes |
|---|---|---|---|---|
| Pre-flight checks | N | 4 | — | — |
| Unit tests | N | N | — | — |
| Assertions (tier) | N | N | N% | — |
| ... | ... | ... | ... | — |
After promptfoo evals complete, launch the interactive web interface:
promptfoo view -y &
# Opens browser at http://localhost:15500
# Shows all evaluations with interactive filtering, pass/fail drill-down, response comparison
# Results are persisted in ~/.promptfoo/promptfoo.db — all historical runs are visible
echo "Promptfoo UI running at http://localhost:15500 — press Ctrl+C to stop"Leave the server running for the user to explore results. The web UI provides:
TODAY=$(date +%Y-%m-%d)
ARCHIVE_DIR="test-reports/binary-test/$TODAY"
mkdir -p "$ARCHIVE_DIR"
# Copy all reports generated during this session
cp test-reports/assertions-report-*.html test-reports/assertions-report-*.jsonl "$ARCHIVE_DIR/" 2>/dev/null
cp test-reports/multi-assertions-report-*.html test-reports/multi-assertions-report-*.jsonl "$ARCHIVE_DIR/" 2>/dev/null
cp test-reports/smart-analysis-*.md "$ARCHIVE_DIR/" 2>/dev/null
cp test-reports/mlx-model-report-*.html test-reports/mlx-model-report-*.jsonl "$ARCHIVE_DIR/" 2>/dev/null
# Copy promptfoo results
PROMPTFOO_DIR="${AFM_PROMPTFOO_OUT_DIR:-/Volumes/edata/promptfoo/data/maclocal-api/current}"
cp "$PROMPTFOO_DIR"/*-mlx-community_*.json "$ARCHIVE_DIR/" 2>/dev/nullWrite a SUMMARY.md in the archive directory with: binary path, version, model tested, platform, date, and a pass/fail table for every test suite run.
| Suite | What it validates |
|---|---|
| Assertions: sections 0-8, 10-15 | Core server functionality |
| Promptfoo: structured, toolcall, grammar (non-concurrent), frameworks | API-level tool calling and structured output |
| OpenAI compat evals | SDK compatibility |
| Batch correctness | KV cache isolation under concurrency |
| Suite | Typical pass rate | Why it varies |
|---|---|---|
| Promptfoo: opencode, pi, openclaw, hermes | 70-90% | Model can't always pick correct tool for complex scenarios |
| Promptfoo: toolcall-quality | ~80% | Model quality on when-to-call decisions |
| Promptfoo: grammar (concurrent) | 50-70% | Known race condition in --concurrent 2 grammar path |
| Smart analysis | Varies | AI judge scoring variance, thinking model token budget |
| Batch multiturn prefix | ~85-90% | Model answer quality at high concurrency |
model_type in config.jsonMLXChatCompletionsController state machine# Smoke test the current build
/test-afm-binary .build/arm64-apple-macosx/release/afm
# Test a Homebrew-installed binary
/test-afm-binary $(brew --prefix afm-next)/bin/afm
# Test a pip-installed binary
/test-afm-binary $(python3 -c "import macafm_next; print(macafm_next.binary_path())")© scouzi1966, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/test-afm-binary of scouzi1966/maclocal-api.
Open the folder on GitHubat commit 138ca5d
Test Afm Binary next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Test Afm Binary this skillscouzi1966/maclocal-api | 345 | — | ~3.8k | Automated safety check: Pass | MIT | |
| Agent Eval Casesagentailor/fullstack-langgraph-nextjs-agent | 132 | — | ~5.3k | Automated safety check: Pass | MIT | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Agents Best PracticesDenisSergeevitch/agents-best-practices | 2.4k | — | ~7.4k | Automated safety check: Pass | MIT | |
| AgentSquad for Swift2FastLabs/agent-squad | 7.8k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 |
agentailor/fullstack-langgraph-nextjs-agent
Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
DenisSergeevitch/agents-best-practices
A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.
2FastLabs/agent-squad
Guides building on-device multi-agent apps in Swift with the AgentSquad framework: which agent, orchestrator, classifier, storage or voice type fits each situation.
raullenchai/Rapid-MLX
Autonomous performance optimization: research, PoC, benchmark, implement, review, PR
scouzi1966/maclocal-api
Maintain and extend AFM (maclocal-api), a Swift OpenAI-compatible local LLM server and CLI for Apple Foundation Models, MLX models, API gateway proxying, and Vision OCR.
scouzi1966/maclocal-api
Build AFM from scratch — submodules, patches, webui, and Swift build.
scouzi1966/maclocal-api
Run and review the Promptfoo-based AFM agentic evaluation suite.
scouzi1966/maclocal-api
A skill your agent uses when user wants to build a PyPI wheel from an existing compiled afm binary and publish to PyPI.
scouzi1966/maclocal-api
Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.
scouzi1966/maclocal-api
A skill your agent uses when testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors, classifies them as afm translation bugs vs model generation errors, and…
Categories
Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…. Test Afm Binary is an agent skill from scouzi1966/maclocal-api. Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU profiling.
Test Afm Binary fits situations like: user wants to validate a binary post-build; after code changes.
Run `npx skills add scouzi1966/maclocal-api --skill test-afm-binary -a claude-code`. Or copy the skill folder (.claude/skills/test-afm-binary in scouzi1966/maclocal-api) into .claude/skills/test-afm-binary in your project. Claude Code loads it when a task matches its description.
Run `npx skills add scouzi1966/maclocal-api --skill test-afm-binary -a codex`. Or copy the skill folder (.claude/skills/test-afm-binary in scouzi1966/maclocal-api) into .agents/skills/test-afm-binary in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scouzi1966/maclocal-api --skill test-afm-binary -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-afm-binary, .gemini/skills/test-afm-binary, .github/skills/test-afm-binary and .opencode/skills/test-afm-binary in your project.
Going by SKILL.md and its folder, Test Afm Binary needs the command-line tools its instructions call (python3, brew and swift). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Test Afm Binary is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Test Afm Binary: Agent Eval Cases (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Agents Best Practices (DenisSergeevitch/agents-best-practices, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
scouzi1966 (a GitHub user) maintains it in scouzi1966/maclocal-api, which has 345 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 5, 2026.
Source: scouzi1966/maclocal-api on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.