Kane CLI Browser Testing
LambdaTest/kane-cli
Drives a real browser through the kane-cli tool and designs requirement-linked test suites from a PRD or a plain description, with mobile and cloud-grid runs.
Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.
$ npx skills add scouzi1966/maclocal-api --skill test-macafm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install scouzi1966/maclocal-api test-macafm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/test-macafm .claude/skills/test-macafm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "test-macafm" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-macafm into .claude/skills/test-macafm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-macafm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-macafmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add scouzi1966/maclocal-api --skill test-macafm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install scouzi1966/maclocal-api test-macafm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/test-macafm .agents/skills/test-macafm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "test-macafm" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-macafm into .agents/skills/test-macafm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-macafm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scouzi1966/maclocal-api --skill test-macafm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install scouzi1966/maclocal-api test-macafm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/test-macafm .cursor/skills/test-macafm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "test-macafm" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-macafm into .cursor/skills/test-macafm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-macafm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/scouzi1966/maclocal-api.git --path .claude/skills/test-macafm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add scouzi1966/maclocal-api --skill test-macafm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install scouzi1966/maclocal-api test-macafm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/test-macafm .gemini/skills/test-macafm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "test-macafm" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-macafm into .gemini/skills/test-macafm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-macafm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install scouzi1966/maclocal-api test-macafmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add scouzi1966/maclocal-api --skill test-macafm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/test-macafm .github/skills/test-macafm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "test-macafm" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-macafm into .github/skills/test-macafm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-macafm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scouzi1966/maclocal-api --skill test-macafm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install scouzi1966/maclocal-api test-macafm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/test-macafm .opencode/skills/test-macafm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "test-macafm" agent skill from https://github.com/scouzi1966/maclocal-api/tree/main/.claude/skills/test-macafm into .opencode/skills/test-macafm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-macafm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
test-macafmRun the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.
Test Macafm is an agent skill from scouzi1966/maclocal-api. Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis. Use when asked to test, validate, regression-check, or benchmark AFM before release, after code changes, or for model onboarding.
Its SKILL.md is about 7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/interpreting-scores.md` and `references/test-inventory.md`).
It sits in AI & LLM Engineering, covering Test generation. It works with macOS. The repository describes itself as: 'afm' command cli: macOS server and single prompt mode that exposes Apple's Foundation and MLX Models and other APIs running on your Mac through a single aggregated…. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 138ca5d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3swiftcurlnpmFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl and npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Test Macafm loads about 7k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 57 tokens; SKILL.md has 2,522 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from scouzi1966/maclocal-api at commit 138ca5d, republished under its MIT licence (© scouzi1966). 2,522 words, ~6,965 tokens.
.claude/skills/test-macafm/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Run the maclocal-api test suite: automated pass/fail assertions and smart analysis (the smart suite's AI judge is opt-in — default off; ask the user before enabling it).
Use this skill when the user asks to:
.build/release/afm. Ask if user has a custom build location.| Tier | Time | When to use | What runs |
|---|---|---|---|
| smoke | ~2 min | Quick sanity check, any small model, CI | test-assertions.sh --tier smoke |
| standard | ~15 min | After feature changes, mid-size model | test-assertions.sh --tier standard |
| full | ~60 min | Release validation, production model | test-assertions.sh --tier full + mlx-model-test.sh (smart suite; AI judge opt-in — ask the user) with test-llm-comprehensive.txt + promptfoo agentic evals |
Quick guide:
# Ensure release build is current
swift build -c releaseMACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache \
.build/release/afm mlx -m MODEL --port 9998 \
--tool-call-parser afm_adaptive_xml \
--enable-prefix-caching \
--enable-grammar-constraints &
# Wait for server to be ready
until curl -sf http://127.0.0.1:9998/v1/models >/dev/null 2>&1; do sleep 1; doneRecommended flags for testing:
--tool-call-parser afm_adaptive_xml — best tool call parser with JSON-in-XML fallback--enable-prefix-caching — 67-79% prompt token savings on repeated requests--enable-grammar-constraints — EBNF constrained decoding forces valid XML tool calls, improving success from 60% to 100% on realistic workloads./Scripts/test-assertions.sh --tier TIER --model MODEL --port 9998Interpret results immediately. If any FAIL, investigate before proceeding.
The smart analysis harness manages its own server (port 9877) — do NOT pass --port.
It uses test-llm-comprehensive.txt which has an [all] baseline prompt and [@ label]
template sections.
The AI judge is OPT-IN — default OFF. By default, run the smart suite WITHOUT an AI
judge: it executes every prompt and records the model's raw outputs to the report for
manual review, with no claude/codex scoring. Before running, ask the user (e.g.
via AskUserQuestion) whether to enable the AI judge — it adds latency/cost and invokes an
external CLI:
"Run the smart suite with an AI judge (claude) scoring each response, or without it (just record outputs for manual review)? Default: without."
Default — no AI judge (records outputs only; omit --smart entirely):
AFM_BIN=.build/release/afm ./Scripts/mlx-model-test.sh \
--model MODEL \
--prompts Scripts/test-llm-comprehensive.txtOnly if the user opts in — append --smart 1:claude (the --smart flag accepts a batch
mode prefix and tool list):
AFM_BIN=.build/release/afm ./Scripts/mlx-model-test.sh \
--model MODEL \
--prompts Scripts/test-llm-comprehensive.txt \
--smart 1:claudeSmart analysis options (only when the AI judge is enabled):
--smart claude or --smart codex — batch mode 0 (one big swoop, may fail on large test suites)--smart 1:claude or --smart 1:codex — batch mode 1 (test-by-test, more reliable)--smart 1:claude,codex — run multiple AI judges--tests 1,5,10 — run only specific test numbers (1-indexed)Note: The [all] prompt runs for every test variant. With high max_tokens (e.g., 32768
on code tests), thinking models may generate very long reasoning for the baseline prompt.
Total run time for full suite: ~45-90 min depending on model speed.
Generates an interactive HTML report with measured DRAM bandwidth, GPU utilization/power timelines, and per-kernel Metal shader names from xctrace Shader Timeline.
One-time setup (creates custom Instruments template with Shader Timeline enabled):
python3 Scripts/create-shader-template.pyRun the profile (no server needed — uses single-prompt mode):
python3 Scripts/gpu-profile-report.py MODEL [max_tokens] [prompt]
# Default: 4096 tokens, built-in GPU analysis prompt
# Example: python3 Scripts/gpu-profile-report.py mlx-community/Qwen3.5-35B-A3B-4bitThis does everything automatically:
--gpu-profile --gpu-trace 15/tmp/afm-gpu-profile.html and opens in browserOr use individual flags on any AFM invocation:
afm mlx -m MODEL --gpu-profile -s "prompt" # Zero-overhead stats
afm mlx -m MODEL --gpu-profile-bw -s "prompt" # + mactop bandwidth (~5s)
afm mlx -m MODEL --gpu-trace 10 -s "prompt" # xctrace shader traceLive bandwidth monitor (run in separate terminal during server requests):
./Scripts/gpu-profile.sh bandwidthWhat the report shows:
What to look for:
affine_qmv_fast (decode bottleneck), steel_gemm_fused (prefill), sdpa_vector (attention)Clients can request GPU profiling data via the X-AFM-Profile HTTP header.
No server flags required — works on any running AFM server.
Two levels:
# Summary: GPU power, memory, bandwidth, tok/s
curl http://127.0.0.1:9999/v1/chat/completions \
-H "X-AFM-Profile: true" \
-d '{"model":"m","messages":[{"role":"user","content":"Hi"}]}'
# Extended: summary + 300ms time-series samples (for charts/dashboards)
curl http://127.0.0.1:9999/v1/chat/completions \
-H "X-AFM-Profile: extended" \
-d '{"model":"m","messages":[{"role":"user","content":"Hi"}]}'Response fields (afm_profile):
gpu_power_avg_w / gpu_power_peak_w — GPU power via native IOReport (no mactop)memory_weights_gib / memory_kv_gib / memory_peak_gib — memory breakdown in GiBprefill_tok_s / decode_tok_s — throughputest_bandwidth_gbs — DRAM bandwidth from IOReport power (calibrated at startup via MLX GPU stress)chip / theoretical_bw_gbs — hardware contextgpu_samples — number of 300ms readings takenExtended adds (afm_profile_extended):
summary — same as afm_profilesamples[] — per-300ms readings: {t, bw_gbs, gpu_pct, gpu_power_w, dram_power_w}How it works internally:
Energy Model + GPU Stats channels sampled every 300ms via DispatchSource timer[DONE]) and non-streamingWhat to look for:
gpu_power_peak_w ~28W during decode on M3 Ultra (matches mactop)est_bandwidth_gbs ~170-180 GB/s for Qwen3.5-35B-A3B-4bit (21% of 800 GB/s theoretical)afm_profile absent from response when header not sent (no null pollution)The promptfoo agentic eval suite tests AFM's tool-calling and structured-output across multiple server configurations and real-world agent framework schemas. It manages its own server lifecycle.
Prerequisites: promptfoo CLI must be installed (npm install -g promptfoo).
Run the full suite:
AFM_MODEL=MODEL \
AFM_BINARY=.build/arm64-apple-macosx/release/afm \
MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache \
./Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh allRun individual suites:
# Just structured output tests
./Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh structured
# Just tool calling (all 3 parser profiles)
./Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh toolcall
# Just grammar constraint validation (8 server phases)
./Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh grammar-constraints
# Just one agent framework
./Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh opencodeAvailable modes: all, structured, structured-stress, toolcall, toolcall-quality, grammar-constraints, agentic, frameworks, opencode, pi, openclaw, hermes, default, adaptive-xml, adaptive-xml-grammar
| Suite | Tests | Profiles | What it validates |
|---|---|---|---|
| structured | 6 | 1 (api json_schema) | response_format=json_schema strict compliance |
| structured-stress | 4 | 1 | Nested arrays, enums, nullable types in schema |
| toolcall | 7 | 3 (default, adaptive-xml, grammar) | Basic tool call parsing: weather, time, multi-tool |
| toolcall-quality | 6 | 3 | BFCL-inspired when-to-call decisions (should model use a tool?) |
| grammar-constraints | 17 | 8 server phases | Schema + tool enforcement across: no-grammar, grammar-enabled, adaptive-xml, concurrent, prefix-cache, mixed-strict, header downgrade/enforce |
| agentic | 4 | 3 | Multi-turn coding workflow tool chains |
| frameworks | 8 | 3 | Agent framework tool shapes (OpenCode, Pi, OpenClaw, Hermes) |
| opencode | 37 | 3 | OpenCode built-in tools (primary-source derived) |
| pi | 20 | 3 | Pi coding-agent tools |
| openclaw | 12 | 3 | OpenClaw tool coverage |
| hermes | 12 | 3 | Hermes agentic framework tools |
| Profile | AFM flags | Purpose |
|---|---|---|
default | (none) | Baseline: auto-detected tool call format |
adaptive-xml | --tool-call-parser afm_adaptive_xml | Adaptive XML with JSON-in-XML fallback |
adaptive-xml-grammar | --tool-call-parser afm_adaptive_xml --enable-grammar-constraints | Adaptive XML + EBNF grammar enforcement |
grammar-enabled | --enable-grammar-constraints | Grammar without adaptive XML |
grammar-enabled-adaptive-xml | Both flags | Regression guard: grammar + adaptive XML |
grammar-enabled-concurrent | --enable-grammar-constraints --concurrent 2 | Grammar under concurrency |
grammar-enabled-prefix-cache | --enable-grammar-constraints --enable-prefix-caching | Grammar + prefix caching interaction |
grammar-enabled-concurrent-cache | All three flags | Full feature stack |
providers/afm_provider.mjs — Custom promptfoo provider with two transports: api (OpenAI-compatible HTTP) and cli-guided-json (direct binary invocation). Supports extract modes: content, tool_calls, normalized_message, full_response. Captures responseHeaders for grammar header assertions.judges/assert-grammar-header.mjs — Validates X-Grammar-Constraints response header: expects "downgraded" when grammar not available, absent when grammar active.judges/classify-failures.mjs — Post-run AI-based failure classifier: categorizes each failure as afm_bug (server/protocol), model_quality (wrong tool/args), or harness_bug (false negative).| Variable | Default | Purpose |
|---|---|---|
AFM_MODEL | mlx-community/Qwen3.5-35B-A3B-4bit | Model to test |
AFM_BINARY | .build/arm64-apple-macosx/release/afm | Binary path |
AFM_PROMPTFOO_OUT_DIR | /Volumes/edata/promptfoo/data/maclocal-api/current | Report output dir |
AFM_PROMPTFOO_PORT | 9999 | Server port |
MACAFM_MLX_MODEL_CACHE | (none) | Model cache dir |
JSON reports per suite+profile in $AFM_PROMPTFOO_OUT_DIR:
structured-MODEL_SLUG.jsontoolcall-{default,adaptive-xml,adaptive-xml-grammar}-MODEL_SLUG.jsongrammar-{schema,tools}-{no-grammar,grammar-enabled,adaptive-xml,concurrent,prefix-cache}-MODEL_SLUG.json{agentic,frameworks,opencode,pi,openclaw,hermes}-{default,adaptive-xml,adaptive-xml-grammar}-MODEL_SLUG.jsontest-reports/assertions-report-*.htmltest-reports/smart-analysis-{tool}-*.mdtest-reports/mlx-model-report-*.html/tmp/afm-gpu-profile.html (+ /tmp/afm-metal.trace for Instruments)test-reports/assertions-report-*.jsonl, test-reports/mlx-model-report-*.jsonl$AFM_PROMPTFOO_OUT_DIR/{suite}-{profile}-MODEL_SLUG.json (default: /Volumes/edata/promptfoo/data/maclocal-api/current/)kill %1 # or whatever the background job is| Group | Common failures | What to check |
|---|---|---|
| Stop | Stop string found in output | Check MLXModelService.swift stop buffer logic, streaming vs non-streaming paths |
| Logprobs | Schema invalid, logprob > 0 | Check resolveLogprobs() and buildChoiceLogprobs() |
| Think | <think> tags in content | Check extractThinkContent() and extractThinkTags() |
| Tools | No tool_calls, invalid JSON args | Check extractToolCallsFallback(), model's tool call format |
| Cache | cached_tokens always 0 | Check enablePrefixCaching, findPrefixLength(), PromptCacheBox |
| Concurrent | Non-200 responses | Check SerialAccessContainer locking, request queuing |
| Error | Wrong HTTP status codes | Check controller validation logic |
| Kwargs | Thinking not disabled by enable_thinking: false | Check chat_template_kwargs merging into additionalContext in MLXModelService.swift |
| Perf | Low tok/s, high TTFT | Check model quantization, Metal kernel performance |
| OpenAI-compat | Stream usage chunk missing, logprobs absent | Check StreamingUsageChunk encoding, empty choices on final chunk |
| Guided JSON | Schema validation failure, invalid JSON | Check --guided-json / response_format pipeline, grammar constraints |
| Batch | Garbage output, wrong answers at B>1 | Check BatchScheduler, KV cache isolation, mask generation |
Known patterns where AI judges score incorrectly (see references/interpreting-scores.md):
<think> — correct, not a bugmax_tokens budget on reasoning with empty visible content — model behavior, not a server bug[all] baseline prompt scored low when it runs with a code/math test's high max_tokens and system prompt — irrelevant context for the baseline prompt| Category | Typical pass rate | What failures mean |
|---|---|---|
| structured, structured-stress | 100% | Server bug in response_format pipeline — investigate immediately |
| toolcall (all profiles) | 100% | Server bug in tool call parsing — investigate immediately |
| toolcall-quality | ~80% | Model chose wrong tool or missed when-to-call — model quality, not server |
| grammar-schema / grammar-tools (non-concurrent) | 100% | Grammar constraint enforcement broken — server bug |
| grammar-schema / grammar-tools (concurrent) | ~50-70% | Known race condition in --concurrent 2 grammar path — not release blocker |
| grammar-header / grammar-mixed | 100% | X-Grammar-Constraints header or mixed-strict wiring broken — server bug |
| agentic | ~75-100% | Multi-turn failures are usually model quality; 0% pass = server bug |
| frameworks | 100% | Framework tool shapes must parse correctly — server bug if failing |
| opencode | ~70-80% | Complex 37-tool scenarios; model can't always pick correct tool — model quality |
| pi | ~80-90% | Model prompt injection resistance varies — model quality |
| openclaw | ~80-85% | Model quality on OpenClaw-specific schemas |
| hermes | ~90-100% | Hermes format failures on adaptive-xml profiles = parser difference, not bug |
Key rule: structured, toolcall, grammar-* (non-concurrent), frameworks suites should be 100% pass. Any failure there is a server bug. Everything else has model-quality variance.
Post-run failure classification (optional): Run judges/classify-failures.mjs on any result JSON to get AI-based afm_bug vs model_quality vs harness_bug classification.
ToolCallFormat.infer() and model's config.jsonScripts/apply-mlx-patches.sh --checkFull-harness concurrency sweep that starts the server, runs warmup, tests all concurrency levels, collects GPU metrics via mactop, saves JSON results, and generates a comparison chart.
Scripts/benchmarks/benchmark_afm_vs_mlxlm.py
# AFM-only concurrency sweep (recommended for quick benchmarks)
python3 Scripts/benchmarks/benchmark_afm_vs_mlxlm.py --afm-only
# Full AFM vs mlx-lm comparison (both servers, fair A/B)
python3 Scripts/benchmarks/benchmark_afm_vs_mlxlm.py
# Re-generate graph from existing results
python3 Scripts/benchmarks/benchmark_afm_vs_mlxlm.py --graph
python3 Scripts/benchmarks/benchmark_afm_vs_mlxlm.py --graph Scripts/benchmark-results/FILE.json--concurrent N[1, 2, 4, 8, 12, 16, 20, 24, 32, 40, 50]Scripts/benchmark-results/concurrency-benchmark-TIMESTAMP.jsonScripts/benchmark-results/concurrency-benchmark-TIMESTAMP.png| Variable | Default | Purpose |
|---|---|---|
MODEL_ID | mlx-community/Qwen3.5-35B-A3B-4bit | Model to benchmark |
MAX_TOKENS | 4096 | Tokens per request (forces long decode) |
MAX_CONCURRENT | 50 | --concurrent flag value (must be >= max level) |
LEVELS | [1,2,4,8,12,16,20,24,32,40,50] | Concurrency levels to test |
AFM_PORT | 9999 | Port for AFM server |
B Agg t/s Per-req Wall GPU% GPU W
1 118.7 118.7 34.5s 94% 28.5W
2 193.9 97.0 42.2s 93% 41.6W
4 298.4 74.6 54.9s 97% 62.7W
8 407.3 50.9 80.5s 96% 75.5W
12 493.4 41.1 99.6s 98% 83.4W
16 573.9 35.9 114.2s 99% 88.2W
20 581.6 29.1 140.8s 98% 79.1W
24 629.6 27.4 149.6s 99% 83.2W| Script | Purpose |
|---|---|
Scripts/feature-mlx-concurrent-batch/batch_stress_mactop.py | Quick stress test at arbitrary concurrency (client-only, needs running server on port 9876) |
Scripts/feature-mlx-concurrent-batch/batch_stress_ioreg.py | Same but uses ioreg for GPU stats (less accurate) |
Scripts/feature-mlx-concurrent-batch/validate_responses.py | Known-answer correctness at B={1,2,4,8} |
Scripts/feature-mlx-concurrent-batch/validate_mixed_workload.py | Mixed short+long workload batch validation |
Scripts/feature-mlx-concurrent-batch/validate_multiturn_prefix.py | Multi-turn prefix cache under concurrency |
| File | Purpose |
|---|---|
Scripts/benchmarks/benchmark_afm_vs_mlxlm.py | Full concurrency benchmark harness (server lifecycle, warmup, sweep, GPU metrics, chart generation) |
Scripts/test-assertions.sh | Automated pass/fail assertion tests (unit/smoke/standard/full tiers, includes swift test) |
Scripts/test-llm-comprehensive.txt | Comprehensive smart analysis test suite (model-generic, [@ label] template mode, has [all] baseline) |
Scripts/test-Qwen3.5-35B-A3B-4bit.txt | Model-specific test suite for Qwen3.5-35B-A3B-4bit (same tests as comprehensive, hardcoded model) |
Scripts/test-edge-cases.txt | Legacy smart analysis test prompts (smaller set) |
Scripts/test-sampling-params.sh | Sampling parameter tests (seed, temp, top_p, etc.) |
Scripts/test-structured-outputs.sh | JSON schema / structured output tests |
Scripts/test-tool-call-parsers.py | Unit tests for tool call parsing |
Scripts/mlx-model-test.sh | Test harness: runs prompts, collects results, generates reports |
Scripts/test-chat-template-kwargs.sh | Standalone chat_template_kwargs tests (includes --no-think CLI + precedence) |
Scripts/regression-test.sh | Quick regression smoke test |
Scripts/feature-codex-optimize-api/test-openai-compat-evals.py | OpenAI-python SDK compatibility evals (non-stream, stream, logprobs, vllm bench) |
Scripts/feature-codex-optimize-api/test-guided-json-evals.py | Guided JSON / structured output evals (API, streaming, CLI, SDK parse, edge cases) |
Scripts/feature-mlx-concurrent-batch/validate_responses.py | Batched generation correctness: known-answer questions at B={1,2,4,8} |
Scripts/feature-mlx-concurrent-batch/validate_mixed_workload.py | Mixed short+long workload batch validation with GPU metrics |
Scripts/feature-mlx-concurrent-batch/validate_multiturn_prefix.py | Multi-turn prefix cache validation under concurrency |
Scripts/gpu-profile-report.py | Full GPU shader profiling harness: mactop BW + --gpu-profile + --gpu-trace + HTML report |
Scripts/gpu-profile.sh | GPU profiling helpers: bandwidth monitor, capture, trace, power |
Scripts/create-shader-template.py | One-time: patches Metal System Trace template for per-kernel shader names |
Tests/MacLocalAPITests/StreamingUsageChunkTests.swift | Unit tests: streaming usage chunks, finish reasons, Foundation commonPrefixLength |
Tests/MacLocalAPITests/ConcurrentBatchTests.swift | Unit tests: RequestSlot, StreamChunk, BatchScheduler internals |
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh | Promptfoo agentic eval orchestrator: 11 modes, 8 server profiles, 16 configs |
Scripts/feature-promptfoo-agentic/providers/afm_provider.mjs | Custom promptfoo provider: api + cli-guided-json transports, 4 extract modes |
Scripts/feature-promptfoo-agentic/judges/assert-grammar-header.mjs | Custom assertion: validates X-Grammar-Constraints response header |
Scripts/feature-promptfoo-agentic/judges/classify-failures.mjs | AI-based failure classifier: afm_bug vs model_quality vs harness_bug |
Scripts/feature-promptfoo-agentic/promptfooconfig.*.yaml | 16 promptfoo config files (~137 test cases total) |
Scripts/feature-promptfoo-agentic/datasets/ | 16 YAML dataset files across structured, toolcall, grammar, agentic directories |
enable_thinking=false disables thinking (if model supports it)test-openai-compat-evals.py (non-stream, stream, logprobs, usage chunk)test-guided-json-evals.py (API schema, streaming schema, SDK parse)--smart 1:claude / --smart 1:codex)validate_responses.py at B={1,2,4,8}validate_mixed_workload.py (short+long decode, GPU metrics)validate_multiturn_prefix.py (multi-turn conversations under concurrency)gpu-profile-report.py (bandwidth, power, kernel names, HTML report)X-AFM-Profile: true returns afm_profile with GPU power + bandwidthX-AFM-Profile: extended returns afm_profile_extended with samples arrayafm_profile fields in response (no null pollution)[DONE]© scouzi1966, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in .claude/skills/test-macafm of scouzi1966/maclocal-api.
Open the folder on GitHubat commit 138ca5d
Test Macafm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Test Macafm this skillscouzi1966/maclocal-api | 346 | — | ~7k | Automated safety check: Pass | MIT | |
| Kane CLI Browser TestingLambdaTest/kane-cli | 248 | — | ~8.4k | Automated safety check: Pass | Apache-2.0 | |
| Synthetic Eval Data Generatorai-evals-course/evals-skills | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Hipfire Testerwarpfront/hipfire | 655 | — | ~1.5k | Automated safety check: Pass | Custom licence | |
| QAteam-attention/hoyeon | 173 | — | ~2.6k | Automated safety check: Notes | MIT | |
| Test Cross Platformr3bl-org/r3bl-open-core | 485 | — | ~771 | Automated safety check: Pass | Apache-2.0 |
LambdaTest/kane-cli
Drives a real browser through the kane-cli tool and designs requirement-linked test suites from a PRD or a plain description, with mobile and cloud-grid runs.
ai-evals-course/evals-skills
Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.
warpfront/hipfire
Guide a tester through hipfire bring-up, serve smoke, claim-scoped harnesses, and benchmark reporting on AMD RDNA/CDNA GPUs.
team-attention/hoyeon
Systematically QA test any application — web apps, native macOS apps, Electron apps, CLI tools, interactive REPLs, or anything on screen.
r3bl-org/r3bl-open-core
Synchronize repository changes to remote test fleet (macOS, Windows) and execute the full test suite across all platforms concurrently.
gustavscirulis/snapgrid
Generate test templates for unit tests, integration tests, and UI tests using Swift Testing and XCTest.
scouzi1966/maclocal-api
Maintain and extend AFM (maclocal-api), a Swift OpenAI-compatible local LLM server and CLI for Apple Foundation Models, MLX models, API gateway proxying, and Vision OCR.
scouzi1966/maclocal-api
Build AFM from scratch — submodules, patches, webui, and Swift build.
scouzi1966/maclocal-api
Run and review the Promptfoo-based AFM agentic evaluation suite.
scouzi1966/maclocal-api
Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…
scouzi1966/maclocal-api
A skill your agent uses when user wants to build a PyPI wheel from an existing compiled afm binary and publish to PyPI.
scouzi1966/maclocal-api
A skill your agent uses when testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors, classifies them as afm translation bugs vs model generation errors, and…
Works with
Categories
Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis. Test Macafm is an agent skill from scouzi1966/maclocal-api. Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.
Test Macafm fits situations like: regression-check; benchmark AFM before release; after code changes; for model onboarding.
Run `npx skills add scouzi1966/maclocal-api --skill test-macafm -a claude-code`. Or copy the skill folder (.claude/skills/test-macafm in scouzi1966/maclocal-api) into .claude/skills/test-macafm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add scouzi1966/maclocal-api --skill test-macafm -a codex`. Or copy the skill folder (.claude/skills/test-macafm in scouzi1966/maclocal-api) into .agents/skills/test-macafm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scouzi1966/maclocal-api --skill test-macafm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-macafm, .gemini/skills/test-macafm, .github/skills/test-macafm and .opencode/skills/test-macafm in your project.
Going by SKILL.md and its folder, Test Macafm needs the command-line tools its instructions call (python3, swift, curl and npm). Our summary lists: Python 3; Node.js.
SKILL.md contains no URLs. Its commands use curl and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Test Macafm is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Test Macafm: Kane CLI Browser Testing (LambdaTest/kane-cli, 248 stars), Synthetic Eval Data Generator (ai-evals-course/evals-skills, 1.5k stars), Hipfire Tester (warpfront/hipfire, 655 stars) and QA (team-attention/hoyeon, 173 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
scouzi1966 (a GitHub user) maintains it in scouzi1966/maclocal-api, which has 346 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 8, 2026.
Source: scouzi1966/maclocal-api on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.