Agent skill

Add Afm Model

by scouzi1966 in scouzi1966/maclocal-api

Add support for a new HuggingFace MLX model to AFM. An agent skill from scouzi1966/maclocal-api.

MITAuto-check passedAI & LLM Engineering

Install Add Afm Model

skills CLI
$ npx skills add scouzi1966/maclocal-api --skill add-afm-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scouzi1966/maclocal-api add-afm-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scouzi1966/maclocal-api.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/add-afm-model .claude/skills/add-afm-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-afm-model
GitHub stars
346
Token cost
~1.4k tokens
SKILL.md length
636 words
Files
3 (incl. references)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Add support for a new HuggingFace MLX model to AFM. An agent skill from scouzi1966/maclocal-api.

  • Works in 6 steps: Parse Input → Fetch config.json → Extract model_type → …
  • User wants to add
  • SKILL.md covers Usage, Instructions, Tier: Registry Addition and Tier: Port from Python, plus 2 more sections
  • Calls swift; reaches huggingface.co and github.com

What it does

Add Afm Model is an agent skill from scouzi1966/maclocal-api. Add support for a new HuggingFace MLX model to AFM. Use when user wants to add, onboard, or check compatibility of a model — handles everything from "already supported" to implementing new architectures.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/implementation-guide.md` and `references/model-investigation.md`).

It sits in AI & LLM Engineering, covering Model hubs and datasets and iOS development. It works with Hugging Face and macOS. The repository describes itself as: 'afm' command cli: macOS server and single prompt mode that exposes Apple's Foundation and MLX Models and other APIs running on your Mac through a single aggregated…. The licence is MIT.

When your agent uses it

  • User wants to add
  • Check compatibility of a model — handles everything from already supported to implementing new architectures

Example prompts

  • “already supported”
  • “/add-afm-model”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Parse Input
  2. Fetch config.json
  3. Extract model_type
  4. Check LLMTypeRegistry
  5. Check Existing Architectures
  6. Check Python mlx-lm

What it can do on your machine

Read from SKILL.md and the folder at commit 138ca5d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • swift

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Afm Model loads about 1.4k tokens when it runs, and up to ~4.6k if it reads all its reference files. Until then it costs about 54 tokens; SKILL.md has 636 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scouzi1966/maclocal-api at commit 138ca5d, republished under its MIT licence (© scouzi1966). 636 words, ~1,396 tokens.

Download SKILL.mdSave it as .claude/skills/add-afm-model/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
add-afm-model
description
Add support for a new HuggingFace MLX model to AFM. Use when user wants to add, onboard, or check compatibility of a model — handles everything from "already supported" to implementing new architectures.
user_invocable
true

Add AFM Model

Investigate and add support for a HuggingFace MLX model to the AFM server.

Usage

  • /add-afm-model <model-id> — e.g., /add-afm-model mlx-community/Qwen3-8B-4bit
  • /add-afm-model <url> — e.g., /add-afm-model https://huggingface.co/mlx-community/Qwen3-8B-4bit

Instructions

Step 1: Parse Input

Extract the HuggingFace model ID from the user's input:

  • Full URL: https://huggingface.co/org/model → org/model
  • Model ID: org/model → use as-is
  • If ambiguous, ask the user.
Step 2: Fetch config.json

Fetch https://huggingface.co/<model-id>/resolve/main/config.json using WebFetch.

Not MLX check: If config.json has no quantization or quantization_config field, the model is not MLX-quantized. Inform the user:

"This model is not in MLX format. Look for an MLX-quantized version on huggingface.co/mlx-community, or quantize it yourself with mlx_lm.convert."

Stop here if not MLX.

Step 3: Extract model_type

Read the model_type field from config.json. This is the key that maps to a Swift model implementation.

Also note: architectures, num_experts/num_local_experts (MoE indicator), image_token_id/vision_config (VLM indicator).

Step 4: Check LLMTypeRegistry

Search Scripts/patches/LLMModelFactory.swift for the model_type string in the LLMTypeRegistry.shared dictionary (lines ~25-80).

If found → Already supported. Tell the user:

"This model is already supported! Run it with:

MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache afm mlx -m <model-id> --port 9999
```"

Suggest running /test-macafm for validation if this is a new model variant.

Stop here if already registered.

Step 5: Check Existing Architectures

The model_type is NOT in the registry. Now determine if an existing Swift implementation can handle it.

  1. List files in vendor/mlx-swift-lm/Libraries/MLXLLM/Models/
  2. Search for the model's base architecture name (e.g., if model_type is foo_moe, check for Foo.swift or similar)
  3. Read the model's config.json fields and compare against existing implementations — some architectures handle variants (e.g., DeepseekV3 handles kimi_k2, Qwen2 handles acereason)
  4. Check if an existing model has a dense fallback (e.g., numExperts == 0 path)

If a compatible architecture exists → Registry-only fix. Proceed to Tier: Registry Addition.

Step 6: Check Python mlx-lm

Search for the model_type in the Python mlx-lm library:

  • Fetch https://github.com/ml-explore/mlx-lm/tree/main/mlx_lm/models (or search GitHub)
  • Look for a Python file matching the model_type

If Python implementation exists → Port to Swift. Proceed to Tier: Port from Python.

If no Python implementation → Implement from scratch. Proceed to Tier: New Architecture.


Tier: Registry Addition

The simplest case — architecture exists, just needs a type alias.

  1. Read references/implementation-guide.md for the patch system details
  2. Add the model_type to LLMTypeRegistry.shared in Scripts/patches/LLMModelFactory.swift, mapping to the correct Configuration/Model pair
  3. If the model is a VLM (has vision_config/image_token_id), also add to Scripts/patches/VLMModelFactory.swift
  4. If the model has a new tool call format, update Scripts/patches/ToolCallFormat.swift infer() method
  5. Apply patches: ./Scripts/apply-mlx-patches.sh
  6. Build: swift build (or /build-afm for full rebuild)
  7. Verify: start server and test with a simple prompt
Show full SKILL.md (214 more words)Show less

Tier: Port from Python

Port the Python mlx-lm implementation to Swift.

  1. Read references/implementation-guide.md for the full implementation pattern
  2. Read references/model-investigation.md for config.json field mapping
  3. Fetch and study the Python implementation from mlx-lm
  4. Find the closest existing Swift model to use as a template (check vendor/mlx-swift-lm/Libraries/MLXLLM/Models/)
  5. Create the new Swift file in Scripts/patches/<ModelName>.swift
  6. Implement: Configuration (Codable) → Attention → MLP → TransformerBlock → Model
  7. Add to PATCH_FILES, TARGET_PATHS, NEW_FILES in Scripts/apply-mlx-patches.sh
  8. Register in LLMTypeRegistry (and VLM if needed)
  9. Add weight sanitization if needed (in the Configuration's sanitize())
  10. Apply, build, verify

Tier: New Architecture

No existing implementation anywhere. Research and implement from scratch.

  1. Read references/implementation-guide.md and references/model-investigation.md
  2. Find architecture documentation: paper, blog post, or reference implementation
  3. Study config.json thoroughly for all architecture-specific fields
  4. Find the closest existing Swift model as a starting template
  5. Follow the same implementation steps as "Port from Python" above
  6. Pay special attention to: attention patterns, normalization layers, MoE routing, positional embeddings
  7. Test extensively — new architectures often have subtle bugs

Key Reminders

  • NEVER edit files in vendor/ directly — all changes go through Scripts/patches/
  • Always check VLM registry too if the model has vision capabilities
  • Use MACAFM_MLX_MODEL_CACHE=/Volumes/edata/models/vesta-test-cache to avoid re-downloading
  • After adding a new model, suggest running /test-macafm for full validation

© scouzi1966, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .claude/skills/add-afm-model of scouzi1966/maclocal-api.

  • SKILL.md
  • references/implementation-guide.md
  • references/model-investigation.md

Open the folder on GitHubat commit 138ca5d

Compare with similar skills

Add Afm Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Afm Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Afm Model this skillscouzi1966/maclocal-api346—~1.4kAutomated safety check: PassMIT
Edge Bringupexeex/edge-cores110—~1.7kAutomated safety check: NotesApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
AgentSquad for Swift2FastLabs/agent-squad7.8k—~3.5kAutomated safety check: PassApache-2.0
Esmfold2JimLiu/science-skills2284 repos~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Edge Bringup

    exeex/edge-cores

    Prepare a macOS or Ubuntu machine for edge-e3 development, diagnose missing Verilator/LLVM/Python dependencies, initialize the public repository, and answer or act on the example prompts in the root…

    110 GitHub stars~1.7k tokensUpdated 14 days ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • AgentSquad for Swift

    2FastLabs/agent-squad

    Guides building on-device multi-agent apps in Swift with the AgentSquad framework: which agent, orchestrator, classifier, storage or voice type fits each situation.

    7.8k GitHub stars~3.5k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    228 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from scouzi1966/maclocal-api

All 12 skills in this repo
  • Afm

    scouzi1966/maclocal-api

    Maintain and extend AFM (maclocal-api), a Swift OpenAI-compatible local LLM server and CLI for Apple Foundation Models, MLX models, API gateway proxying, and Vision OCR.

    346 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Build Afm

    scouzi1966/maclocal-api

    Build AFM from scratch — submodules, patches, webui, and Swift build.

    346 GitHub stars~1.8k tokensUpdated today
    Auto-check: notes
  • Codex Promptfoo Agentic Eval

    scouzi1966/maclocal-api

    Run and review the Promptfoo-based AFM agentic evaluation suite.

    346 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Test Afm Binary

    scouzi1966/maclocal-api

    Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…

    346 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Afm Release Wheel

    scouzi1966/maclocal-api

    A skill your agent uses when user wants to build a PyPI wheel from an existing compiled afm binary and publish to PyPI.

    346 GitHub stars~1.5k tokensUpdated today
    Auto-check: warnings
  • Test Macafm

    scouzi1966/maclocal-api

    Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.

    346 GitHub stars~7k tokensUpdated today
    Auto-check passed

Questions about Add Afm Model

What does Add Afm Model do?

Add support for a new HuggingFace MLX model to AFM. An agent skill from scouzi1966/maclocal-api. Add Afm Model is an agent skill from scouzi1966/maclocal-api. Add support for a new HuggingFace MLX model to AFM.

When should I use Add Afm Model?

Add Afm Model fits situations like: user wants to add; check compatibility of a model — handles everything from already supported to implementing new architectures.

How do I install Add Afm Model in Claude Code?

Run `npx skills add scouzi1966/maclocal-api --skill add-afm-model -a claude-code`. Or copy the skill folder (.claude/skills/add-afm-model in scouzi1966/maclocal-api) into .claude/skills/add-afm-model in your project. Claude Code loads it when a task matches its description.

How do I install Add Afm Model in Codex?

Run `npx skills add scouzi1966/maclocal-api --skill add-afm-model -a codex`. Or copy the skill folder (.claude/skills/add-afm-model in scouzi1966/maclocal-api) into .agents/skills/add-afm-model in your project. Codex loads it when a task matches its description.

Can I use Add Afm Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scouzi1966/maclocal-api --skill add-afm-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-afm-model, .gemini/skills/add-afm-model, .github/skills/add-afm-model and .opencode/skills/add-afm-model in your project.

What does Add Afm Model need to run?

Going by SKILL.md and its folder, Add Afm Model needs the command-line tools its instructions call (swift). Our summary lists: Python 3.

Does Add Afm Model access the network?

SKILL.md names 2 domains. In commands or code: huggingface.co and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Add Afm Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Afm Model use?

Add Afm Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Afm Model use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.2k tokens, read only when the agent opens those files.

What are the alternatives to Add Afm Model?

Skills that share tags, products or a category with Add Afm Model: Edge Bringup (exeex/edge-cores, 110 stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and AgentSquad for Swift (2FastLabs/agent-squad, 7.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Afm Model?

scouzi1966 (a GitHub user) maintains it in scouzi1966/maclocal-api, which has 346 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 10, 2026.

Source: scouzi1966/maclocal-api on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.