Agent skill

Test Macafm

by scouzi1966 in scouzi1966/maclocal-api

Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.

MITAuto-check passedAI & LLM Engineering

Install Test Macafm

skills CLI
$ npx skills add scouzi1966/maclocal-api --skill test-macafm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scouzi1966/maclocal-api test-macafm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/test-macafm .claude/skills/test-macafm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-macafm
GitHub stars
346
Token cost
~7k tokens
SKILL.md length
2,522 words
Files
3 (incl. references)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.

  • Works in 8 steps: Build Check → Start Server (if not running) → Run Automated Assertions → …
  • Regression-check
  • SKILL.md covers Triggers, First Questions to Ask, Tier Decision Tree and Execution Workflow, plus 4 more sections
  • Calls python3, swift and curl

What it does

Test Macafm is an agent skill from scouzi1966/maclocal-api. Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis. Use when asked to test, validate, regression-check, or benchmark AFM before release, after code changes, or for model onboarding.

Its SKILL.md is about 7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/interpreting-scores.md` and `references/test-inventory.md`).

It sits in AI & LLM Engineering, covering Test generation. It works with macOS. The repository describes itself as: 'afm' command cli: macOS server and single prompt mode that exposes Apple's Foundation and MLX Models and other APIs running on your Mac through a single aggregated…. The licence is MIT.

When your agent uses it

  • Regression-check
  • Benchmark AFM before release
  • After code changes
  • For model onboarding

Example prompts

  • “/test-macafm”

Requirements

  • Python 3
  • Node.js

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Build Check
  2. Start Server (if not running)
  3. Run Automated Assertions
  4. Run Smart Analysis (full tier only)
  5. Run GPU Shader Profile (full tier, or when investigating perf)
  6. Run Promptfoo Agentic Evals (full tier, or when validating tool calling / structured output)
  7. Review Reports
  8. Stop Server (if we started it)

What it can do on your machine

Read from SKILL.md and the folder at commit 138ca5d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • swift
    • curl
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Macafm loads about 7k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 57 tokens; SKILL.md has 2,522 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scouzi1966/maclocal-api at commit 138ca5d, republished under its MIT licence (© scouzi1966). 2,522 words, ~6,965 tokens.

Download SKILL.mdSave it as .claude/skills/test-macafm/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
test-macafm
description
Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis. Use when asked to test, validate, regression-check, or benchmark AFM before release, after code changes, or for model onboarding.

test-macafm

Run the maclocal-api test suite: automated pass/fail assertions and smart analysis (the smart suite's AI judge is opt-in — default off; ask the user before enabling it).

Triggers

Use this skill when the user asks to:

  • Test or validate the server (e.g., "run the tests", "test AFM", "validate the build")
  • Regression check after code changes
  • Onboard a new model (verify it works correctly with the server)
  • Release check before tagging or pushing
  • Benchmark or profile model performance

First Questions to Ask

  1. Model — Which model to test? (Ask if not specified. Default: whatever's loaded.)
  2. Tier — smoke / standard / full? (Suggest based on context.)
  3. Binary path — Default .build/release/afm. Ask if user has a custom build location.
  4. Port — Default 9998. Ask if user's server is on a different port.
  5. Server running? — Is the server already running, or should tests start it?

Tier Decision Tree

TierTimeWhen to useWhat runs
smoke~2 minQuick sanity check, any small model, CItest-assertions.sh --tier smoke
standard~15 minAfter feature changes, mid-size modeltest-assertions.sh --tier standard
full~60 minRelease validation, production modeltest-assertions.sh --tier full + mlx-model-test.sh (smart suite; AI judge opt-in — ask the user) with test-llm-comprehensive.txt + promptfoo agentic evals

Quick guide:

  • "Just run a quick test" → smoke
  • "Test before merging" → standard
  • "Full release validation" or "onboard new model" → full
  • User doesn't specify → suggest standard

Execution Workflow

1. Build Check
bash
# Ensure release build is current
swift build -c release
2. Start Server (if not running)
bash
MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache \
  .build/release/afm mlx -m MODEL --port 9998 \
  --tool-call-parser afm_adaptive_xml \
  --enable-prefix-caching \
  --enable-grammar-constraints &
# Wait for server to be ready
until curl -sf http://127.0.0.1:9998/v1/models >/dev/null 2>&1; do sleep 1; done

Recommended flags for testing:

  • --tool-call-parser afm_adaptive_xml — best tool call parser with JSON-in-XML fallback
  • --enable-prefix-caching — 67-79% prompt token savings on repeated requests
  • --enable-grammar-constraints — EBNF constrained decoding forces valid XML tool calls, improving success from 60% to 100% on realistic workloads
3. Run Automated Assertions
bash
./Scripts/test-assertions.sh --tier TIER --model MODEL --port 9998

Interpret results immediately. If any FAIL, investigate before proceeding.

4. Run Smart Analysis (full tier only)

The smart analysis harness manages its own server (port 9877) — do NOT pass --port. It uses test-llm-comprehensive.txt which has an [all] baseline prompt and [@ label] template sections.

The AI judge is OPT-IN — default OFF. By default, run the smart suite WITHOUT an AI judge: it executes every prompt and records the model's raw outputs to the report for manual review, with no claude/codex scoring. Before running, ask the user (e.g. via AskUserQuestion) whether to enable the AI judge — it adds latency/cost and invokes an external CLI:

"Run the smart suite with an AI judge (claude) scoring each response, or without it (just record outputs for manual review)? Default: without."

Default — no AI judge (records outputs only; omit --smart entirely):

bash
AFM_BIN=.build/release/afm ./Scripts/mlx-model-test.sh \
  --model MODEL \
  --prompts Scripts/test-llm-comprehensive.txt

Only if the user opts in — append --smart 1:claude (the --smart flag accepts a batch mode prefix and tool list):

bash
AFM_BIN=.build/release/afm ./Scripts/mlx-model-test.sh \
  --model MODEL \
  --prompts Scripts/test-llm-comprehensive.txt \
  --smart 1:claude

Smart analysis options (only when the AI judge is enabled):

  • --smart claude or --smart codex — batch mode 0 (one big swoop, may fail on large test suites)
  • --smart 1:claude or --smart 1:codex — batch mode 1 (test-by-test, more reliable)
  • --smart 1:claude,codex — run multiple AI judges
  • --tests 1,5,10 — run only specific test numbers (1-indexed)

Note: The [all] prompt runs for every test variant. With high max_tokens (e.g., 32768 on code tests), thinking models may generate very long reasoning for the baseline prompt. Total run time for full suite: ~45-90 min depending on model speed.

5. Run GPU Shader Profile (full tier, or when investigating perf)

Generates an interactive HTML report with measured DRAM bandwidth, GPU utilization/power timelines, and per-kernel Metal shader names from xctrace Shader Timeline.

One-time setup (creates custom Instruments template with Shader Timeline enabled):

bash
python3 Scripts/create-shader-template.py

Run the profile (no server needed — uses single-prompt mode):

bash
python3 Scripts/gpu-profile-report.py MODEL [max_tokens] [prompt]
# Default: 4096 tokens, built-in GPU analysis prompt
# Example: python3 Scripts/gpu-profile-report.py mlx-community/Qwen3.5-35B-A3B-4bit

This does everything automatically:

  1. Warms up mactop (bandwidth monitor, no sudo)
  2. Runs inference with --gpu-profile --gpu-trace 15
  3. Collects 300ms bandwidth/GPU/power samples via PTY during inference
  4. Extracts shader kernel names from the xctrace trace
  5. Generates /tmp/afm-gpu-profile.html and opens in browser

Or use individual flags on any AFM invocation:

bash
afm mlx -m MODEL --gpu-profile -s "prompt"           # Zero-overhead stats
afm mlx -m MODEL --gpu-profile-bw -s "prompt"        # + mactop bandwidth (~5s)
afm mlx -m MODEL --gpu-trace 10 -s "prompt"          # xctrace shader trace

Live bandwidth monitor (run in separate terminal during server requests):

bash
./Scripts/gpu-profile.sh bandwidth

What the report shows:

  • Device info (chip, memory, architecture)
  • Prefill/decode tok/s with exact timing
  • Memory breakdown (model weights vs KV cache)
  • DRAM bandwidth timeline chart (measured via mactop)
  • GPU utilization & power timeline chart
  • Per-kernel Metal shader names (from Shader Timeline)
  • Exact command line for reproducibility

What to look for:

  • GPU utilization <100% during decode → CPU-GPU pipeline bubbles
  • Bandwidth utilization >80% → memory-bound, kernel optimization won't help
  • Bandwidth utilization <20% with MoE model → normal (only active experts read)
  • Key kernels: affine_qmv_fast (decode bottleneck), steel_gemm_fused (prefill), sdpa_vector (attention)
5b. API-Based GPU Profiling (per-request, no CLI flags needed)

Clients can request GPU profiling data via the X-AFM-Profile HTTP header. No server flags required — works on any running AFM server.

Two levels:

bash
# Summary: GPU power, memory, bandwidth, tok/s
curl http://127.0.0.1:9999/v1/chat/completions \
  -H "X-AFM-Profile: true" \
  -d '{"model":"m","messages":[{"role":"user","content":"Hi"}]}'

# Extended: summary + 300ms time-series samples (for charts/dashboards)
curl http://127.0.0.1:9999/v1/chat/completions \
  -H "X-AFM-Profile: extended" \
  -d '{"model":"m","messages":[{"role":"user","content":"Hi"}]}'

Response fields (afm_profile):

  • gpu_power_avg_w / gpu_power_peak_w — GPU power via native IOReport (no mactop)
  • memory_weights_gib / memory_kv_gib / memory_peak_gib — memory breakdown in GiB
  • prefill_tok_s / decode_tok_s — throughput
  • est_bandwidth_gbs — DRAM bandwidth from IOReport power (calibrated at startup via MLX GPU stress)
  • chip / theoretical_bw_gbs — hardware context
  • gpu_samples — number of 300ms readings taken

Extended adds (afm_profile_extended):

  • summary — same as afm_profile
  • samples[] — per-300ms readings: {t, bw_gbs, gpu_pct, gpu_power_w, dram_power_w}

How it works internally:

  • IOReport Energy Model + GPU Stats channels sampled every 300ms via DispatchSource timer
  • DRAM bandwidth derived from DRAM power using chip-specific calibration constant
  • Calibration runs once at startup: 1 GiB MLX GPU stress test (~2s, async, non-blocking)
  • Per-request isolation: concurrent profiled requests are guarded (second request skips gracefully)
  • Zero overhead when header not sent (one string lookup per request)
  • Works for both streaming (SSE event before [DONE]) and non-streaming

What to look for:

  • gpu_power_peak_w ~28W during decode on M3 Ultra (matches mactop)
  • est_bandwidth_gbs ~170-180 GB/s for Qwen3.5-35B-A3B-4bit (21% of 800 GB/s theoretical)
  • Short requests (<300ms): at least 1 sample (timer first-fires at 100ms)
  • afm_profile absent from response when header not sent (no null pollution)
6. Run Promptfoo Agentic Evals (full tier, or when validating tool calling / structured output)

The promptfoo agentic eval suite tests AFM's tool-calling and structured-output across multiple server configurations and real-world agent framework schemas. It manages its own server lifecycle.

Prerequisites: promptfoo CLI must be installed (npm install -g promptfoo).

Run the full suite:

bash
AFM_MODEL=MODEL \
AFM_BINARY=.build/arm64-apple-macosx/release/afm \
MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache \
./Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh all

Run individual suites:

bash
# Just structured output tests
./Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh structured

# Just tool calling (all 3 parser profiles)
./Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh toolcall

# Just grammar constraint validation (8 server phases)
./Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh grammar-constraints

# Just one agent framework
./Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.sh opencode

Available modes: all, structured, structured-stress, toolcall, toolcall-quality, grammar-constraints, agentic, frameworks, opencode, pi, openclaw, hermes, default, adaptive-xml, adaptive-xml-grammar

Suite Coverage (~137 test cases across 16 configs)
SuiteTestsProfilesWhat it validates
structured61 (api json_schema)response_format=json_schema strict compliance
structured-stress41Nested arrays, enums, nullable types in schema
toolcall73 (default, adaptive-xml, grammar)Basic tool call parsing: weather, time, multi-tool
toolcall-quality63BFCL-inspired when-to-call decisions (should model use a tool?)
grammar-constraints178 server phasesSchema + tool enforcement across: no-grammar, grammar-enabled, adaptive-xml, concurrent, prefix-cache, mixed-strict, header downgrade/enforce
agentic43Multi-turn coding workflow tool chains
frameworks83Agent framework tool shapes (OpenCode, Pi, OpenClaw, Hermes)
opencode373OpenCode built-in tools (primary-source derived)
pi203Pi coding-agent tools
openclaw123OpenClaw tool coverage
hermes123Hermes agentic framework tools
Server Profiles (managed automatically by the script)
ProfileAFM flagsPurpose
default(none)Baseline: auto-detected tool call format
adaptive-xml--tool-call-parser afm_adaptive_xmlAdaptive XML with JSON-in-XML fallback
adaptive-xml-grammar--tool-call-parser afm_adaptive_xml --enable-grammar-constraintsAdaptive XML + EBNF grammar enforcement
grammar-enabled--enable-grammar-constraintsGrammar without adaptive XML
grammar-enabled-adaptive-xmlBoth flagsRegression guard: grammar + adaptive XML
grammar-enabled-concurrent--enable-grammar-constraints --concurrent 2Grammar under concurrency
grammar-enabled-prefix-cache--enable-grammar-constraints --enable-prefix-cachingGrammar + prefix caching interaction
grammar-enabled-concurrent-cacheAll three flagsFull feature stack
Custom Provider & Judges
  • providers/afm_provider.mjs — Custom promptfoo provider with two transports: api (OpenAI-compatible HTTP) and cli-guided-json (direct binary invocation). Supports extract modes: content, tool_calls, normalized_message, full_response. Captures responseHeaders for grammar header assertions.
  • judges/assert-grammar-header.mjs — Validates X-Grammar-Constraints response header: expects "downgraded" when grammar not available, absent when grammar active.
  • judges/classify-failures.mjs — Post-run AI-based failure classifier: categorizes each failure as afm_bug (server/protocol), model_quality (wrong tool/args), or harness_bug (false negative).
Environment Variables
VariableDefaultPurpose
AFM_MODELmlx-community/Qwen3.5-35B-A3B-4bitModel to test
AFM_BINARY.build/arm64-apple-macosx/release/afmBinary path
AFM_PROMPTFOO_OUT_DIR/Volumes/edata/promptfoo/data/maclocal-api/currentReport output dir
AFM_PROMPTFOO_PORT9999Server port
MACAFM_MLX_MODEL_CACHE(none)Model cache dir
Output

JSON reports per suite+profile in $AFM_PROMPTFOO_OUT_DIR:

  • structured-MODEL_SLUG.json
  • toolcall-{default,adaptive-xml,adaptive-xml-grammar}-MODEL_SLUG.json
  • grammar-{schema,tools}-{no-grammar,grammar-enabled,adaptive-xml,concurrent,prefix-cache}-MODEL_SLUG.json
  • {agentic,frameworks,opencode,pi,openclaw,hermes}-{default,adaptive-xml,adaptive-xml-grammar}-MODEL_SLUG.json
7. Review Reports
  • Assertion report: test-reports/assertions-report-*.html
  • Smart analysis: test-reports/smart-analysis-{tool}-*.md
  • HTML report: test-reports/mlx-model-report-*.html
  • GPU profile: /tmp/afm-gpu-profile.html (+ /tmp/afm-metal.trace for Instruments)
  • JSONL data: test-reports/assertions-report-*.jsonl, test-reports/mlx-model-report-*.jsonl
  • Promptfoo evals: $AFM_PROMPTFOO_OUT_DIR/{suite}-{profile}-MODEL_SLUG.json (default: /Volumes/edata/promptfoo/data/maclocal-api/current/)
7. Stop Server (if we started it)
bash
kill %1  # or whatever the background job is

Interpreting Results

Assertion Test Failures
GroupCommon failuresWhat to check
StopStop string found in outputCheck MLXModelService.swift stop buffer logic, streaming vs non-streaming paths
LogprobsSchema invalid, logprob > 0Check resolveLogprobs() and buildChoiceLogprobs()
Think<think> tags in contentCheck extractThinkContent() and extractThinkTags()
ToolsNo tool_calls, invalid JSON argsCheck extractToolCallsFallback(), model's tool call format
Cachecached_tokens always 0Check enablePrefixCaching, findPrefixLength(), PromptCacheBox
ConcurrentNon-200 responsesCheck SerialAccessContainer locking, request queuing
ErrorWrong HTTP status codesCheck controller validation logic
KwargsThinking not disabled by enable_thinking: falseCheck chat_template_kwargs merging into additionalContext in MLXModelService.swift
PerfLow tok/s, high TTFTCheck model quantization, Metal kernel performance
OpenAI-compatStream usage chunk missing, logprobs absentCheck StreamingUsageChunk encoding, empty choices on final chunk
Guided JSONSchema validation failure, invalid JSONCheck --guided-json / response_format pipeline, grammar constraints
BatchGarbage output, wrong answers at B>1Check BatchScheduler, KV cache isolation, mask generation
Show full SKILL.md (1,042 more words)Show less
Smart Analysis False Positives

Known patterns where AI judges score incorrectly (see references/interpreting-scores.md):

  • Stop sequences truncating output scored as "low quality" — truncation IS the expected behavior
  • Empty content when stop fires on first visible token — correct behavior
  • JSON mode not constraining thinking models — prompt injection, not grammar-constrained
  • "Missing reasoning" when model doesn't support <think> — correct, not a bug
  • Thinking model consuming entire max_tokens budget on reasoning with empty visible content — model behavior, not a server bug
  • [all] baseline prompt scored low when it runs with a code/math test's high max_tokens and system prompt — irrelevant context for the baseline prompt
Promptfoo Eval Failures
CategoryTypical pass rateWhat failures mean
structured, structured-stress100%Server bug in response_format pipeline — investigate immediately
toolcall (all profiles)100%Server bug in tool call parsing — investigate immediately
toolcall-quality~80%Model chose wrong tool or missed when-to-call — model quality, not server
grammar-schema / grammar-tools (non-concurrent)100%Grammar constraint enforcement broken — server bug
grammar-schema / grammar-tools (concurrent)~50-70%Known race condition in --concurrent 2 grammar path — not release blocker
grammar-header / grammar-mixed100%X-Grammar-Constraints header or mixed-strict wiring broken — server bug
agentic~75-100%Multi-turn failures are usually model quality; 0% pass = server bug
frameworks100%Framework tool shapes must parse correctly — server bug if failing
opencode~70-80%Complex 37-tool scenarios; model can't always pick correct tool — model quality
pi~80-90%Model prompt injection resistance varies — model quality
openclaw~80-85%Model quality on OpenClaw-specific schemas
hermes~90-100%Hermes format failures on adaptive-xml profiles = parser difference, not bug

Key rule: structured, toolcall, grammar-* (non-concurrent), frameworks suites should be 100% pass. Any failure there is a server bug. Everything else has model-quality variance.

Post-run failure classification (optional): Run judges/classify-failures.mjs on any result JSON to get AI-based afm_bug vs model_quality vs harness_bug classification.

When to Escalate
  • SDPA regression: NaN or garbage in long-context tests → check MLX version, see MEMORY.md
  • Tool call format mismatch: Unknown format → check ToolCallFormat.infer() and model's config.json
  • Build failure: Vendor patch conflict → run Scripts/apply-mlx-patches.sh --check

Concurrency Benchmark

Full-harness concurrency sweep that starts the server, runs warmup, tests all concurrency levels, collects GPU metrics via mactop, saves JSON results, and generates a comparison chart.

Script

Scripts/benchmarks/benchmark_afm_vs_mlxlm.py

Usage
bash
# AFM-only concurrency sweep (recommended for quick benchmarks)
python3 Scripts/benchmarks/benchmark_afm_vs_mlxlm.py --afm-only

# Full AFM vs mlx-lm comparison (both servers, fair A/B)
python3 Scripts/benchmarks/benchmark_afm_vs_mlxlm.py

# Re-generate graph from existing results
python3 Scripts/benchmarks/benchmark_afm_vs_mlxlm.py --graph
python3 Scripts/benchmarks/benchmark_afm_vs_mlxlm.py --graph Scripts/benchmark-results/FILE.json
What it does
  1. Detects hardware (chip, memory)
  2. Starts server(s) with --concurrent N
  3. 60s GPU settle + multi-round warmup (JIT kernel compilation)
  4. Sweeps concurrency levels: [1, 2, 4, 8, 12, 16, 20, 24, 32, 40, 50]
  5. At each level: fires N simultaneous streaming 4096-token requests, measures aggregate tok/s, per-request tok/s, GPU power/temp/usage via mactop
  6. Saves JSON to Scripts/benchmark-results/concurrency-benchmark-TIMESTAMP.json
  7. Generates PNG chart to Scripts/benchmark-results/concurrency-benchmark-TIMESTAMP.png
Configuration (top of script)
VariableDefaultPurpose
MODEL_IDmlx-community/Qwen3.5-35B-A3B-4bitModel to benchmark
MAX_TOKENS4096Tokens per request (forces long decode)
MAX_CONCURRENT50--concurrent flag value (must be >= max level)
LEVELS[1,2,4,8,12,16,20,24,32,40,50]Concurrency levels to test
AFM_PORT9999Port for AFM server
Reference results (March 18, v0.9.7, M3 Ultra 512GB, --concurrent 28)
  B   Agg t/s   Per-req   Wall    GPU%   GPU W
  1     118.7     118.7   34.5s    94%   28.5W
  2     193.9      97.0   42.2s    93%   41.6W
  4     298.4      74.6   54.9s    97%   62.7W
  8     407.3      50.9   80.5s    96%   75.5W
 12     493.4      41.1   99.6s    98%   83.4W
 16     573.9      35.9  114.2s    99%   88.2W
 20     581.6      29.1  140.8s    98%   79.1W
 24     629.6      27.4  149.6s    99%   83.2W
Additional batch validation scripts
ScriptPurpose
Scripts/feature-mlx-concurrent-batch/batch_stress_mactop.pyQuick stress test at arbitrary concurrency (client-only, needs running server on port 9876)
Scripts/feature-mlx-concurrent-batch/batch_stress_ioreg.pySame but uses ioreg for GPU stats (less accurate)
Scripts/feature-mlx-concurrent-batch/validate_responses.pyKnown-answer correctness at B={1,2,4,8}
Scripts/feature-mlx-concurrent-batch/validate_mixed_workload.pyMixed short+long workload batch validation
Scripts/feature-mlx-concurrent-batch/validate_multiturn_prefix.pyMulti-turn prefix cache under concurrency

Key File Reference

FilePurpose
Scripts/benchmarks/benchmark_afm_vs_mlxlm.pyFull concurrency benchmark harness (server lifecycle, warmup, sweep, GPU metrics, chart generation)
Scripts/test-assertions.shAutomated pass/fail assertion tests (unit/smoke/standard/full tiers, includes swift test)
Scripts/test-llm-comprehensive.txtComprehensive smart analysis test suite (model-generic, [@ label] template mode, has [all] baseline)
Scripts/test-Qwen3.5-35B-A3B-4bit.txtModel-specific test suite for Qwen3.5-35B-A3B-4bit (same tests as comprehensive, hardcoded model)
Scripts/test-edge-cases.txtLegacy smart analysis test prompts (smaller set)
Scripts/test-sampling-params.shSampling parameter tests (seed, temp, top_p, etc.)
Scripts/test-structured-outputs.shJSON schema / structured output tests
Scripts/test-tool-call-parsers.pyUnit tests for tool call parsing
Scripts/mlx-model-test.shTest harness: runs prompts, collects results, generates reports
Scripts/test-chat-template-kwargs.shStandalone chat_template_kwargs tests (includes --no-think CLI + precedence)
Scripts/regression-test.shQuick regression smoke test
Scripts/feature-codex-optimize-api/test-openai-compat-evals.pyOpenAI-python SDK compatibility evals (non-stream, stream, logprobs, vllm bench)
Scripts/feature-codex-optimize-api/test-guided-json-evals.pyGuided JSON / structured output evals (API, streaming, CLI, SDK parse, edge cases)
Scripts/feature-mlx-concurrent-batch/validate_responses.pyBatched generation correctness: known-answer questions at B={1,2,4,8}
Scripts/feature-mlx-concurrent-batch/validate_mixed_workload.pyMixed short+long workload batch validation with GPU metrics
Scripts/feature-mlx-concurrent-batch/validate_multiturn_prefix.pyMulti-turn prefix cache validation under concurrency
Scripts/gpu-profile-report.pyFull GPU shader profiling harness: mactop BW + --gpu-profile + --gpu-trace + HTML report
Scripts/gpu-profile.shGPU profiling helpers: bandwidth monitor, capture, trace, power
Scripts/create-shader-template.pyOne-time: patches Metal System Trace template for per-kernel shader names
Tests/MacLocalAPITests/StreamingUsageChunkTests.swiftUnit tests: streaming usage chunks, finish reasons, Foundation commonPrefixLength
Tests/MacLocalAPITests/ConcurrentBatchTests.swiftUnit tests: RequestSlot, StreamChunk, BatchScheduler internals
Scripts/feature-promptfoo-agentic/run-promptfoo-agentic.shPromptfoo agentic eval orchestrator: 11 modes, 8 server profiles, 16 configs
Scripts/feature-promptfoo-agentic/providers/afm_provider.mjsCustom promptfoo provider: api + cli-guided-json transports, 4 extract modes
Scripts/feature-promptfoo-agentic/judges/assert-grammar-header.mjsCustom assertion: validates X-Grammar-Constraints response header
Scripts/feature-promptfoo-agentic/judges/classify-failures.mjsAI-based failure classifier: afm_bug vs model_quality vs harness_bug
Scripts/feature-promptfoo-agentic/promptfooconfig.*.yaml16 promptfoo config files (~137 test cases total)
Scripts/feature-promptfoo-agentic/datasets/16 YAML dataset files across structured, toolcall, grammar, agentic directories

Validation Checklist

Smoke Tier
  • Server reachable, model loaded
  • Basic completion returns content
  • Stop sequences work (absent from output, correct finish_reason)
  • Logprobs schema valid
  • Think extraction works (if model supports it)
  • Basic tool call works
  • Error handling (empty messages, malformed JSON)
Standard Tier (adds)
  • All smoke checks
  • Streaming stop sequence parity
  • Streaming logprobs
  • Prompt cache: cached_tokens=0 first, >0 second
  • Concurrent requests (2 and 3 simultaneous)
  • Multi-tool calls
  • Additional stop edge cases
  • chat_template_kwargs: enable_thinking=false disables thinking (if model supports it)
  • chat_template_kwargs: streaming parity
  • chat_template_kwargs: default behavior unaffected
  • OpenAI-compat evals: test-openai-compat-evals.py (non-stream, stream, logprobs, usage chunk)
  • Guided JSON evals: test-guided-json-evals.py (API schema, streaming schema, SDK parse)
Full Tier (adds)
  • All standard checks
  • Performance: TTFT < 5s, tok/s > 1
  • Long context (2K, 4K tokens) no crash/NaN
  • Smart analysis: test-llm-comprehensive.txt (AI judge opt-in — default off, ask the user; enable with --smart 1:claude / --smart 1:codex)
  • Streaming parity (assembled content matches non-streaming)
  • Cache timing improvement visible
  • Batch correctness: validate_responses.py at B={1,2,4,8}
  • Batch mixed workload: validate_mixed_workload.py (short+long decode, GPU metrics)
  • Batch prefix cache: validate_multiturn_prefix.py (multi-turn conversations under concurrency)
  • GPU shader profile: gpu-profile-report.py (bandwidth, power, kernel names, HTML report)
  • API profile: X-AFM-Profile: true returns afm_profile with GPU power + bandwidth
  • API profile: X-AFM-Profile: extended returns afm_profile_extended with samples array
  • API profile: no header → no afm_profile fields in response (no null pollution)
  • API profile: streaming → profile SSE event before [DONE]
  • API profile: concurrent profiled requests → second skips gracefully
  • Promptfoo structured: 100% pass (json_schema + stress)
  • Promptfoo toolcall: 100% pass (all 3 profiles)
  • Promptfoo grammar-constraints (non-concurrent): 100% pass
  • Promptfoo frameworks: 100% pass (all 3 profiles)
  • Promptfoo opencode/pi/openclaw/hermes: >70% pass (model quality variance expected)
  • Promptfoo grammar-header: downgrade/enforce headers correct

© scouzi1966, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .claude/skills/test-macafm of scouzi1966/maclocal-api.

  • SKILL.md
  • references/interpreting-scores.md
  • references/test-inventory.md

Open the folder on GitHubat commit 138ca5d

Compare with similar skills

Test Macafm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Macafm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Macafm this skillscouzi1966/maclocal-api346—~7kAutomated safety check: PassMIT
Kane CLI Browser TestingLambdaTest/kane-cli248—~8.4kAutomated safety check: PassApache-2.0
Synthetic Eval Data Generatorai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.0
Hipfire Testerwarpfront/hipfire655—~1.5kAutomated safety check: PassCustom licence
QAteam-attention/hoyeon173—~2.6kAutomated safety check: NotesMIT
Test Cross Platformr3bl-org/r3bl-open-core485—~771Automated safety check: PassApache-2.0

Similar skills

  • Kane CLI Browser Testing

    LambdaTest/kane-cli

    Drives a real browser through the kane-cli tool and designs requirement-linked test suites from a PRD or a plain description, with mobile and cloud-grid runs.

    248 GitHub stars~8.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Synthetic Eval Data Generator

    ai-evals-course/evals-skills

    Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

    1.5k GitHub stars~1.4k tokensUpdated 14 days ago
    AI & LLM EngineeringAuto-check passed
  • Hipfire Tester

    warpfront/hipfire

    Guide a tester through hipfire bring-up, serve smoke, claim-scoped harnesses, and benchmark reporting on AMD RDNA/CDNA GPUs.

    655 GitHub stars~1.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • QA

    team-attention/hoyeon

    Systematically QA test any application — web apps, native macOS apps, Electron apps, CLI tools, interactive REPLs, or anything on screen.

    173 GitHub stars~2.6k tokensUpdated 4 mo ago
    Productivity & AutomationAuto-check: notes
  • Test Cross Platform

    r3bl-org/r3bl-open-core

    Synchronize repository changes to remote test fleet (macOS, Windows) and execute the full test suite across all platforms concurrently.

    485 GitHub stars~771 tokensUpdated today
    Testing & QAAuto-check passed
  • Test Generator

    gustavscirulis/snapgrid

    Generate test templates for unit tests, integration tests, and UI tests using Swift Testing and XCTest.

    117 GitHub starsUsed in 1 repo~2.7k tokens
    Testing & QAAuto-check: notes

More from scouzi1966/maclocal-api

All 12 skills in this repo
  • Afm

    scouzi1966/maclocal-api

    Maintain and extend AFM (maclocal-api), a Swift OpenAI-compatible local LLM server and CLI for Apple Foundation Models, MLX models, API gateway proxying, and Vision OCR.

    346 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Build Afm

    scouzi1966/maclocal-api

    Build AFM from scratch — submodules, patches, webui, and Swift build.

    346 GitHub stars~1.8k tokensUpdated today
    Auto-check: notes
  • Codex Promptfoo Agentic Eval

    scouzi1966/maclocal-api

    Run and review the Promptfoo-based AFM agentic evaluation suite.

    346 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Test Afm Binary

    scouzi1966/maclocal-api

    Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…

    346 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Afm Release Wheel

    scouzi1966/maclocal-api

    A skill your agent uses when user wants to build a PyPI wheel from an existing compiled afm binary and publish to PyPI.

    346 GitHub stars~1.5k tokensUpdated today
    Auto-check: warnings
  • Test Opencode Tooling

    scouzi1966/maclocal-api

    A skill your agent uses when testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors, classifies them as afm translation bugs vs model generation errors, and…

    346 GitHub stars~4.2k tokensUpdated today
    Auto-check passed

Works with

Questions about Test Macafm

What does Test Macafm do?

Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis. Test Macafm is an agent skill from scouzi1966/maclocal-api. Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.

When should I use Test Macafm?

Test Macafm fits situations like: regression-check; benchmark AFM before release; after code changes; for model onboarding.

How do I install Test Macafm in Claude Code?

Run `npx skills add scouzi1966/maclocal-api --skill test-macafm -a claude-code`. Or copy the skill folder (.claude/skills/test-macafm in scouzi1966/maclocal-api) into .claude/skills/test-macafm in your project. Claude Code loads it when a task matches its description.

How do I install Test Macafm in Codex?

Run `npx skills add scouzi1966/maclocal-api --skill test-macafm -a codex`. Or copy the skill folder (.claude/skills/test-macafm in scouzi1966/maclocal-api) into .agents/skills/test-macafm in your project. Codex loads it when a task matches its description.

Can I use Test Macafm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scouzi1966/maclocal-api --skill test-macafm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-macafm, .gemini/skills/test-macafm, .github/skills/test-macafm and .opencode/skills/test-macafm in your project.

What does Test Macafm need to run?

Going by SKILL.md and its folder, Test Macafm needs the command-line tools its instructions call (python3, swift, curl and npm). Our summary lists: Python 3; Node.js.

Does Test Macafm access the network?

SKILL.md contains no URLs. Its commands use curl and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Macafm safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Macafm use?

Test Macafm is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Macafm use?

About 7k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.5k tokens, read only when the agent opens those files.

What are the alternatives to Test Macafm?

Skills that share tags, products or a category with Test Macafm: Kane CLI Browser Testing (LambdaTest/kane-cli, 248 stars), Synthetic Eval Data Generator (ai-evals-course/evals-skills, 1.5k stars), Hipfire Tester (warpfront/hipfire, 655 stars) and QA (team-attention/hoyeon, 173 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Macafm?

scouzi1966 (a GitHub user) maintains it in scouzi1966/maclocal-api, which has 346 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 8, 2026.

Source: scouzi1966/maclocal-api on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.