Agent skill

Test Opencode Tooling

by scouzi1966 in scouzi1966/maclocal-api

A skill your agent uses when testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors, classifies them as afm translation bugs vs model generation errors, and…

MITAuto-check passedAI & LLM Engineering

Install Test Opencode Tooling

skills CLI
$ npx skills add scouzi1966/maclocal-api --skill test-opencode-tooling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scouzi1966/maclocal-api test-opencode-tooling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/test-opencode-tooling .claude/skills/test-opencode-tooling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-opencode-tooling
GitHub stars
345
Token cost
~4.2k tokens
SKILL.md length
1,402 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors, classifies them as afm translation bugs vs model generation errors, and…

  • Works in 7 steps: Setup → Start OpenCode Serve → For Each Model → …
  • Testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors
  • SKILL.md covers When to Use, First Questions to Ask, OpenCode CLI Gotchas and OpenCode Log & Error Data, plus 3 more sections
  • Calls opencode, sqlite3 and git; reaches opencode.ai

What it does

Test Opencode Tooling is an agent skill from scouzi1966/maclocal-api. Use when testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors, classifies them as afm translation bugs vs model generation errors, and produces a diagnostic report without fixing anything

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Translation. It works with OpenAI and macOS. The repository describes itself as: 'afm' command cli: macOS server and single prompt mode that exposes Apple's Foundation and MLX Models and other APIs running on your Mac through a single aggregated…. The licence is MIT.

When your agent uses it

  • Testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors
  • Classifies them as afm translation bugs vs model generation errors
  • Produces a diagnostic report without fixing anything

Example prompts

  • “/test-opencode-tooling”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Setup
  2. Start OpenCode Serve
  3. For Each Model
  4. Stop OpenCode Serve
  5. Analyze Logs
  6. Generate Report
  7. Present Results

What it can do on your machine

Read from SKILL.md and the folder at commit 138ca5d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • opencode
    • sqlite3
    • git
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • opencode.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Opencode Tooling loads about 4.2k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 1,402 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scouzi1966/maclocal-api at commit 138ca5d, republished under its MIT licence (© scouzi1966). 1,402 words, ~4,180 tokens.

Download SKILL.mdSave it as .claude/skills/test-opencode-tooling/SKILL.md (or your agent's skills folder).
name
test-opencode-tooling
description
Use when testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors, classifies them as afm translation bugs vs model generation errors, and produces a diagnostic report without fixing anything

Test OpenCode Tooling

Automated loop that runs OpenCode tasks against afm, captures tool call errors from both sides, classifies each as an afm bug or model error, and generates a report. Does not fix anything.

When to Use

  • After changing tool call parsing code (XML, streaming, type coercion)
  • Onboarding a new model to verify tool call reliability
  • Investigating user-reported tool call failures with OpenCode
  • Comparing tool call error rates across models

First Questions to Ask

  1. Prompt/PRD — Ask the user to paste the prompt text or provide a file path. This is the task OpenCode will execute (e.g., a PRD, coding task, or test scenario that exercises tool calls).
  2. Model(s) — Which model(s) to test? Show available:
    bash
    MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache ./Scripts/list-models.sh
  3. afm start parameters — Any extra flags beyond defaults? (e.g., --tool-call-parser afm_adaptive_xml, --enable-prefix-caching, --enable-grammar-constraints, --no-think). Recommended: --tool-call-parser afm_adaptive_xml --enable-grammar-constraints — this combination gives the highest tool call success rate (100% on 35B-A3B vs 60% without grammar constraints on realistic workloads).
  4. Iterations — How many times to run the same prompt per model? Default: 1. More runs help distinguish flaky model errors from deterministic afm bugs.
  5. Working directory — Temp dir for OpenCode to work in. Default: create a fresh /tmp/opencode-test-TIMESTAMP per run.

OpenCode CLI Gotchas

CRITICAL: opencode run hangs silently without a PTY. It prints one INFO line and freezes — no error, no output. You must use one of these approaches:

  1. opencode serve + run --attach (recommended): Start a headless server, then attach run to it via expect for PTY
  2. expect wrapper: Provides the pseudo-TTY that opencode run requires

Other gotchas:

  • opencode.json model field must be a string, not an object — "model": "ollama/model-id" not "model": {"default": "..."}
  • The npm provider format (@ai-sdk/openai-compatible) is required for custom baseURL — the "api": "openai" format does NOT accept baseURL
  • OpenCode config is loaded from both ~/.config/opencode/opencode.json (global) AND $WORKDIR/opencode.json (local) — local overrides global
  • The workdir should be a git repo (git init) for OpenCode to function properly

OpenCode Log & Error Data

Log Files (limited — no tool call errors)

OpenCode writes logs to ~/.local/share/opencode/log/ in UTC-timestamped files (e.g., 2026-03-09T172212.log). These logs do NOT contain tool call errors or tool input/output. They only log permission checks, bus events, and registry start/complete.

bash
# Find the latest OpenCode log
ls -t ~/.local/share/opencode/log/*.log | head -1

# Monitor the latest log in real-time
tail -f "$(ls -t ~/.local/share/opencode/log/*.log | head -1)"

Gotcha: Log filenames use UTC timestamps but ls -lt shows local time. A file named 2026-03-10T001322.log was created at 8:13 PM EDT. Use lsof -p <PID> | grep log to find the current session's log file if it doesn't appear in directory listings yet (OpenCode buffers writes).

When monitoring both afm and OpenCode simultaneously:

  • afm log: /tmp/afm-opencode-test.log (or wherever you tee'd it)
  • OpenCode log: ~/.local/share/opencode/log/<latest>.log

ALWAYS start OpenCode with --log-level "DEBUG" --print-logs — both opencode serve and opencode run commands must include these flags.

SQLite Database (structured tool call data with errors)

Tool call inputs, outputs, and errors are stored in OpenCode's SQLite database — not in the log files. This is the only place to get the full JSON of failed tool calls.

Database path: ~/.local/share/opencode/opencode.db

Schema: Tool calls are in the part table as JSON in the data column, keyed by session_id.

bash
# List recent sessions
sqlite3 ~/.local/share/opencode/opencode.db \
  "SELECT id, title, datetime(time_created/1000, 'unixepoch', 'localtime') FROM session ORDER BY time_created DESC LIMIT 5;"

# Get ALL tool call errors for a session
sqlite3 ~/.local/share/opencode/opencode.db \
  "SELECT data FROM part WHERE session_id = '<SESSION_ID>' AND data LIKE '%\"status\":\"error\"%';"

# Get errors for the most recent session
sqlite3 ~/.local/share/opencode/opencode.db \
  "SELECT data FROM part WHERE session_id = (SELECT id FROM session ORDER BY time_created DESC LIMIT 1) AND data LIKE '%\"status\":\"error\"%';"

# Get all edit tool errors across all sessions
sqlite3 ~/.local/share/opencode/opencode.db \
  "SELECT data FROM part WHERE data LIKE '%\"tool\":\"edit\"%' AND data LIKE '%\"status\":\"error\"%' ORDER BY time_created DESC LIMIT 10;"

Error JSON format:

json
{
  "type": "tool",
  "callID": "call_8B05B790A94F4A0EBF2850C0",
  "tool": "edit",
  "state": {
    "status": "error",
    "input": {
      "filePath": "/path/to/file.py",
      "oldString": "text the model expected to find",
      "newString": "replacement text"
    },
    "error": "Error: Could not find oldString in the file. It must match exactly, including whitespace, indentation, and line endings.",
    "time": {
      "start": 1773102597849,
      "end": 1773102597850
    }
  }
}

Successful tool call JSON format:

json
{
  "type": "tool",
  "callID": "call_34CB225B0D184310BD64A839",
  "tool": "edit",
  "state": {
    "status": "completed",
    "input": {
      "filePath": "/path/to/file.py",
      "oldString": "...",
      "newString": "..."
    },
    "output": "Edit applied successfully.",
    "title": "path/to/file.py",
    "metadata": {
      "diagnostics": {},
      "diff": "Index: /path/to/file.py\n===...",
      "filediff": { "file": "...", "before": "...", "after": "..." }
    }
  }
}

Useful queries for test analysis:

bash
# Count tool calls by status for a session
sqlite3 ~/.local/share/opencode/opencode.db \
  "SELECT json_extract(data, '$.tool') as tool,
          json_extract(data, '$.state.status') as status,
          COUNT(*) as cnt
   FROM part
   WHERE session_id = '<SESSION_ID>' AND json_extract(data, '$.type') = 'tool'
   GROUP BY tool, status;"

# Get all tool call inputs/outputs (pipe to jq for pretty-printing)
sqlite3 ~/.local/share/opencode/opencode.db \
  "SELECT data FROM part WHERE session_id = '<SESSION_ID>' AND json_extract(data, '$.type') = 'tool';" | python3 -mjson.tool

Execution Workflow

1. Setup
bash
TIMESTAMP=$(date +%Y%m%d_%H%M%S)
TEST_PORT=9877
OC_PORT=4096
REPORT_DIR="test-reports/opencode-tooling-${TIMESTAMP}"
mkdir -p "$REPORT_DIR"

Save the user's prompt to a file:

bash
cat > "$REPORT_DIR/prompt.md" << 'PROMPT_EOF'
<paste user's prompt here>
PROMPT_EOF
2. Start OpenCode Serve

Create a workdir with git init and config pointing at afm:

bash
OC_WORKDIR="/tmp/opencode-serve-${TIMESTAMP}"
mkdir -p "$OC_WORKDIR"
cd "$OC_WORKDIR" && git init -q && cd -

Write the OpenCode config. Must use npm provider with options.baseURL:

bash
cat > "$OC_WORKDIR/opencode.json" << EOF
{
  "\$schema": "https://opencode.ai/config.json",
  "provider": {
    "ollama": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "afm-test",
      "options": {
        "baseURL": "http://localhost:${TEST_PORT}/v1"
      },
      "models": {
        "${MODEL}": {
          "name": "${MODEL}"
        }
      }
    }
  }
}
EOF

Start the headless server:

bash
cd "$OC_WORKDIR"
opencode serve --port $OC_PORT --print-logs --log-level DEBUG \
  > "$REPORT_DIR/opencode-serve.log" 2>&1 &
OC_SERVE_PID=$!
cd -

# Wait for serve to be ready
until curl -sf http://127.0.0.1:${OC_PORT}/ >/dev/null 2>&1; do sleep 1; done
3. For Each Model
3a. Start afm with verbose logging
bash
AFM_DEBUG=1 MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache \
  .build/release/afm mlx -m "$MODEL" --port $TEST_PORT -V \
  $EXTRA_AFM_FLAGS \
  > "$REPORT_DIR/${MODEL_SLUG}-afm.log" 2>&1 &
AFM_PID=$!

# Wait for server ready
until curl -sf http://127.0.0.1:${TEST_PORT}/v1/models >/dev/null 2>&1; do sleep 1; done

Where MODEL_SLUG is the model ID with / replaced by _.

3b. Run OpenCode via expect + attach (per iteration)

expect provides the PTY that opencode run requires. The --attach flag connects to the serve instance which already has the config and workdir.

bash
/usr/bin/expect << EXPECT_EOF > "$REPORT_DIR/${MODEL_SLUG}-run${RUN}-opencode.json" 2>&1
set timeout 600
log_user 1

spawn opencode run --attach http://localhost:${OC_PORT} --log-level "DEBUG" --print-logs --format json "${PROMPT}"
expect {
    timeout { puts "TIMEOUT"; exit 1 }
    eof { puts "EOF"; exit 0 }
}
EXPECT_EOF

The --format json flag outputs structured JSON events:

  • {"type":"tool_use",...} — tool call with input/output/error
  • {"type":"text",...} — assistant text content
  • {"type":"step_start",...} / {"type":"step_finish",...} — generation boundaries

Each run creates a new session on the same serve instance. The timeout (600s = 10 min) should be enough for most PRDs — increase for complex tasks.

IMPORTANT: Clean workdir between iterations. Before each run, remove all generated files from the OpenCode workdir so that results from a previous iteration don't contaminate the next one (e.g., OpenCode's "must read file before overwriting" guard triggers on leftover files). The cleanest approach is to stop opencode serve, recreate the workdir from scratch (rm -rf "$OC_WORKDIR" && mkdir -p "$OC_WORKDIR" && cd "$OC_WORKDIR" && git init -q && cd -), copy the opencode.json config back, and restart opencode serve. This ensures each iteration starts with a pristine empty git repo.

bash
# Between iterations: reset workdir
kill $OC_SERVE_PID 2>/dev/null; wait $OC_SERVE_PID 2>/dev/null
rm -rf "$OC_WORKDIR"
mkdir -p "$OC_WORKDIR"
cd "$OC_WORKDIR" && git init -q && cd -
# Re-copy opencode.json config (same as setup step)
cat > "$OC_WORKDIR/opencode.json" << EOF
{ ... same config as before ... }
EOF
cd "$OC_WORKDIR"
opencode serve --port $OC_PORT --print-logs --log-level DEBUG \
  >> "$REPORT_DIR/opencode-serve.log" 2>&1 &
OC_SERVE_PID=$!
cd -
until curl -sf http://127.0.0.1:${OC_PORT}/ >/dev/null 2>&1; do sleep 1; done
3c. Stop afm (after all iterations for this model)
bash
kill $AFM_PID 2>/dev/null
wait $AFM_PID 2>/dev/null
4. Stop OpenCode Serve
bash
kill $OC_SERVE_PID 2>/dev/null
3. Analyze Logs

For each run, analyze both log files to extract and classify errors.

Show full SKILL.md (637 more words)Show less
From afm logs (-afm.log), look for:
PatternClassification
SKIP false </tool_call> end tagafm handled correctly (model emitted premature end tag)
EMIT param[N]: key→... with wrong valueCheck if model sent wrong value (model error) or afm mangled it (afm bug)
RECV </tool_call> with raw= bodyRaw model output — compare against what OpenCode received
extractToolCallsFallback activatedIncremental parser failed, fallback used — note if result was correct
SEND tool_call fallback: found 0 tool callsCritical — tool call body couldn't be parsed at all. Usually means model emitted JSON instead of XML inside <tool_call> tags
SEND tool_call name: with JSON in nameafm extracted JSON payload as function name — model mixed formats
coerceArgumentTypes log entriesType coercion activated — check if result matches schema
Malformed XML in raw body (e.g., <function=X> instead of <parameter=X>)Model error — wrong XML tag
Duplicate <parameter=key> tagsModel error — model emitted same param twice
Missing </function> in bodyModel error — incomplete XML generation
From OpenCode output (-opencode.json), look for:
PatternClassification
"tool":"invalid" with mangled tool nameafm parsed function name wrong — cross-ref afm SEND tool_call name: log
"invalid arguments" with undefined valuesParameter was lost — cross-reference afm log to determine if afm dropped it or model never sent it
"expected number, received string"Type coercion failed — afm bug if schema had type: "integer"
Tool name not in schemaModel hallucinated tool — model error
"command" undefined for bash toolCross-ref afm raw body: if <parameter=command> present → afm bug; if <function=command> → model error
SyntaxError with \\\" in written filesPossible afm double-escaping of quotes in tool call arguments
Cross-referencing (the key step):

For each OpenCode error:

  1. Find the corresponding tool call in afm's log (match by timestamp proximity)
  2. Read the raw= body from afm's RECV </tool_call> log
  3. Compare what the model generated vs what afm emitted vs what OpenCode received
  4. Classify:
    • afm schema→model bug: afm sent wrong/incomplete tool schema to the model
    • afm model→client bug: Model output was correct but afm mangled it (dropped param, wrong type, truncated body)
    • Model generation error: Model produced invalid XML, wrong tags, missing params, hallucinated tools
4. Generate Report

Create $REPORT_DIR/report.md:

markdown
# OpenCode Tooling Test Report
- Date: TIMESTAMP
- Model(s): ...
- Prompt: (first 200 chars)
- afm flags: ...
- Iterations per model: N

## Summary
| Model | Runs | Tool Calls | Errors | afm Bugs | Model Errors |
|-------|------|------------|--------|----------|--------------|

## Errors by Category

### afm Translation Bugs (model→client)
| # | Model | Run | Tool | Parameter | What Happened | afm Raw Body |
|---|-------|-----|------|-----------|---------------|-------------|

### afm Translation Bugs (schema→model)
| # | Model | Run | Tool | What Happened |
|---|-------|-----|------|---------------|

### Model Generation Errors
| # | Model | Run | Tool | Error Type | Raw Output |
|---|-------|-----|------|------------|------------|

## Raw Logs
- afm: [link to log file]
- OpenCode: [link to json file]
5. Present Results

Show the user:

  • Summary table (pass rate per model)
  • Each error with classification and evidence
  • Recommendation: which errors are actionable afm bugs vs model limitations
  • Do NOT propose or implement fixes — report only

Error Classification Guide

Definitely afm Bug
  • Parameter present in raw model output but missing in OpenCode's received arguments
  • Type mismatch when schema has explicit type and afm didn't coerce
  • Tool call body truncated (false end tag not caught)
  • Function name mangled or lost
Definitely Model Error
  • JSON inside XML tags (most common): Model emits <tool_call>{"name":"write","arguments":{...}}</tool_call> instead of XML <function=write><parameter=...> format. afm's fallback logs found 0 tool calls — content is silently lost. Qwen3-Coder-Next switches formats unpredictably, especially in longer conversations.
  • <function=X> used instead of <parameter=X> (wrong XML tag)
  • Tool name not in provided schema (hallucinated tool)
  • Parameter never appears in raw model output
  • Garbage characters in parameter values (e.g., trailing })
  • Incomplete XML (missing </function> or </parameter>)
  • <parameter=KEY> without wrapping <function=NAME> — parameters emitted without function context
Ambiguous (needs investigation)
  • Empty parameter value — could be model sending empty or afm dropping content
  • Duplicate parameters — model may emit twice, afm may deduplicate wrong
  • Streaming assembly errors — compare raw chunks vs assembled result
  • Escaped triple quotes (\\\"\\\"\\\") in written files — could be afm double-escaping or model pre-escaping

Common Mistakes

  • Not checking raw afm body: Always cross-reference OpenCode errors against afm's raw= log. Without this, you can't classify.
  • Blaming afm for model errors: Models frequently emit broken XML. Check the raw output first.
  • Blaming the model for afm bugs: afm has had bugs dropping empty params, false end tags, type coercion failures. Don't assume the model is always wrong.
  • Running without -V flag: Without verbose logging, you can't see raw model output or per-parameter emissions. Always use -V.

© scouzi1966, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/test-opencode-tooling of scouzi1966/maclocal-api.

Open the folder on GitHubat commit 138ca5d

Compare with similar skills

Test Opencode Tooling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Opencode Tooling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Opencode Tooling this skillscouzi1966/maclocal-api345—~4.2kAutomated safety check: PassMIT
Pulse Releasequnqin24/Pulse516—~1.3kAutomated safety check: PassApache-2.0
Natural Languagedpearson2699/swift-ios-skills1.2k1 repos~3.5kAutomated safety check: PassCustom licence
Prompt AdaptAgriciDaniel/claude-prompts110—~793Automated safety check: PassMIT
AgentSquad for Swift2FastLabs/agent-squad7.8k—~3.5kAutomated safety check: PassApache-2.0
Perfupraullenchai/Rapid-MLX3.9k—~1.6kAutomated safety check: NotesCustom licence

Similar skills

  • Pulse Release

    qunqin24/Pulse

    Release a new Pulse version end to end — checks, bilingual CHANGELOG entry, VERSION, tag, the release workflow, syncing main, and the issue replies that go with it.

    516 GitHub stars~1.3k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Natural Language

    dpearson2699/swift-ios-skills

    Tokenize, tag, and analyze natural language text using Apple's NaturalLanguage framework and translate between languages with the Translation framework.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Prompt Adapt

    AgriciDaniel/claude-prompts

    Adapt and convert AI prompts between different models and platforms.

    110 GitHub stars~793 tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • AgentSquad for Swift

    2FastLabs/agent-squad

    Guides building on-device multi-agent apps in Swift with the AgentSquad framework: which agent, orchestrator, classifier, storage or voice type fits each situation.

    7.8k GitHub stars~3.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Perfup

    raullenchai/Rapid-MLX

    Autonomous performance optimization: research, PoC, benchmark, implement, review, PR

    3.9k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Mlx Serve

    ddalcu/mlx-serve

    Hook an app, game or script up to the local mlx-serve server for LLM chat, embeddings, image, speech, music, sound effect, video and 3D generation, and Laya/Kev/Clef typed decisions.

    1.8k GitHub stars~1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from scouzi1966/maclocal-api

All 12 skills in this repo
  • Afm

    scouzi1966/maclocal-api

    Maintain and extend AFM (maclocal-api), a Swift OpenAI-compatible local LLM server and CLI for Apple Foundation Models, MLX models, API gateway proxying, and Vision OCR.

    345 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • Build Afm

    scouzi1966/maclocal-api

    Build AFM from scratch — submodules, patches, webui, and Swift build.

    345 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check: notes
  • Codex Promptfoo Agentic Eval

    scouzi1966/maclocal-api

    Run and review the Promptfoo-based AFM agentic evaluation suite.

    345 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • Test Afm Binary

    scouzi1966/maclocal-api

    Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…

    345 GitHub stars~3.8k tokensUpdated 2 days ago
    Auto-check passed
  • Afm Release Wheel

    scouzi1966/maclocal-api

    A skill your agent uses when user wants to build a PyPI wheel from an existing compiled afm binary and publish to PyPI.

    345 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check: warnings
  • Test Macafm

    scouzi1966/maclocal-api

    Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.

    345 GitHub stars~7k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Test Opencode Tooling

What does Test Opencode Tooling do?

A skill your agent uses when testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors, classifies them as afm translation bugs vs model generation errors, and…. Test Opencode Tooling is an agent skill from scouzi1966/maclocal-api.

When should I use Test Opencode Tooling?

Test Opencode Tooling fits situations like: testing tool call reliability between OpenCode and afm — captures streaming XML tool call errors; classifies them as afm translation bugs vs model generation errors; produces a diagnostic report without fixing anything.

How do I install Test Opencode Tooling in Claude Code?

Run `npx skills add scouzi1966/maclocal-api --skill test-opencode-tooling -a claude-code`. Or copy the skill folder (.claude/skills/test-opencode-tooling in scouzi1966/maclocal-api) into .claude/skills/test-opencode-tooling in your project. Claude Code loads it when a task matches its description.

How do I install Test Opencode Tooling in Codex?

Run `npx skills add scouzi1966/maclocal-api --skill test-opencode-tooling -a codex`. Or copy the skill folder (.claude/skills/test-opencode-tooling in scouzi1966/maclocal-api) into .agents/skills/test-opencode-tooling in your project. Codex loads it when a task matches its description.

Can I use Test Opencode Tooling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scouzi1966/maclocal-api --skill test-opencode-tooling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-opencode-tooling, .gemini/skills/test-opencode-tooling, .github/skills/test-opencode-tooling and .opencode/skills/test-opencode-tooling in your project.

What does Test Opencode Tooling need to run?

Going by SKILL.md and its folder, Test Opencode Tooling needs the command-line tools its instructions call (opencode, sqlite3, git and python3). Our summary lists: Python 3.

Does Test Opencode Tooling access the network?

SKILL.md names 1 domain. In commands or code: opencode.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Test Opencode Tooling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Opencode Tooling use?

Test Opencode Tooling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Opencode Tooling use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Opencode Tooling?

Skills that share tags, products or a category with Test Opencode Tooling: Pulse Release (qunqin24/Pulse, 516 stars), Natural Language (dpearson2699/swift-ios-skills, 1.2k stars), Prompt Adapt (AgriciDaniel/claude-prompts, 110 stars) and AgentSquad for Swift (2FastLabs/agent-squad, 7.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Opencode Tooling?

scouzi1966 (a GitHub user) maintains it in scouzi1966/maclocal-api, which has 345 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 5, 2026.

Source: scouzi1966/maclocal-api on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.