Agent skill

Engine Mlx Lm

by autonomous-ai in autonomous-ai/openharness

Serve an MLX or Hugging Face safetensors model already on this Mac with mlx-lm's server and join it to the person's fleet.

MITAuto-check passedAI & LLM Engineering

Install Engine Mlx Lm

skills CLI
$ npx skills add autonomous-ai/openharness --skill engine-mlx-lm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/openharness engine-mlx-lm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/store/agents/autonomous-grid/skills/engine-mlx-lm .claude/skills/engine-mlx-lm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
engine-mlx-lm
GitHub stars
1.1k
Token cost
~1.9k tokens
SKILL.md length
966 words
Files
1
Skills in repo
100
Repo updated
First seen
Licence
MIT

At a glance

Serve an MLX or Hugging Face safetensors model already on this Mac with mlx-lm's server and join it to the person's fleet.

  • Works in 2 steps: The PID is alive. 2. GET /health →… → GET /v1/models → 200 [doc]. 4. One…
  • Tasks that involve Model hubs and datasets
  • SKILL.md covers When to use it, Where its models live / how to…, Installed? Running? and Start an already-downloaded…, plus 6 more sections
  • Calls uv, pip and conda

What it does

Engine Mlx Lm is an agent skill from autonomous-ai/openharness. Serve an MLX or Hugging Face safetensors model already on this Mac with mlx-lm's server and join it to the person's fleet. Load before installing, starting or stopping mlx-lm, or when fleet models lists an mlx or safetensors model on Apple silicon.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model hubs and datasets. It works with Hugging Face and llama.cpp. The repository describes itself as: The ultimate harness for coding agents and beyond. All your agents. All your machines. One command center. Start with code, then follow your curiosity and build across… The licence is MIT.

When your agent uses it

  • Tasks that involve Model hubs and datasets

Example prompts

  • “s server and join it to the person”
  • “/engine-mlx-lm”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. The PID is alive. 2. GET /health → {"status": "ok"} [run] (in the source, not in SERVER.md [doc]).
  2. GET /v1/models → 200 [doc]. 4. One bounded /v1/chat/completions request, max_tokens 16, model

What it can do on your machine

Read from SKILL.md and the folder at commit 50da5db. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • pip
    • conda

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Engine Mlx Lm loads about 1.9k tokens when it runs. Until then it costs about 67 tokens; SKILL.md has 966 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from autonomous-ai/openharness at commit 50da5db, republished under its MIT licence (© autonomous-ai). 966 words, ~1,888 tokens.

Download SKILL.mdSave it as .claude/skills/engine-mlx-lm/SKILL.md (or your agent's skills folder).
name
engine-mlx-lm
description
Serve an MLX or Hugging Face safetensors model already on this Mac with mlx-lm's server and join it to the person's fleet. Load before installing, starting or stopping mlx-lm, or when `fleet models` lists an `mlx` or `safetensors` model on Apple silicon.

mlx-lm server (Apple silicon)

Official docs, read 2026-09-29: SERVER.md · server.py (flags, routes) · README · Hugging Face cache variables · Hugging Face cache layout Tested: mlx-lm 0.31.3 (the latest on PyPI that day), uv venv with Python 3.12, MacBook Pro M1 Pro 32 GB, macOS 26.6, 2026-09-29. Tags: [doc] official source above, [run] seen on the tested machine, [?] unverified.

When to use it

  • fleet models lists a model with format: mlx (an MLX-converted folder: config.json with a quantization block) on a Mac: this is its engine. Grid's own engine reads only GGUF.
  • format: safetensors (a plain Hugging Face model) on a Mac: mlx-lm served one directly [run]; vLLM and SGLang are for NVIDIA/AMD servers.
  • If LM Studio is installed and already holds the MLX model, LM Studio can serve it with nothing new installed (engine-lm-studio). If a GGUF of the same model exists, prefer Grid's engine.
  • The docs call the server "not recommended for production" (basic security only) [doc]: bind 127.0.0.1.
  • A GGUF of the same model already on disk is served by Grid's engine; do not download an MLX copy of it.
  • mlx-lm installed and no MLX model yet: "$GRID_FLEET" candidates mlx (see run-local-model).
  • Verified end to end [run]: started offline on a local snapshot, fleet verify passed ready, answer, tool call and speed; grid join … --at http://127.0.0.1:P/v1 -m <snapshot path> --advertise-as NAME joined in 6 s and the relay answered through the grid.

Where its models live / how to list them

  • Hugging Face cache: $HF_HUB_CACHE, else $HF_HOME/hub, else ~/.cache/huggingface/hub [doc]. Layout models--<org>--<name>/snapshots/<commit>/ with files linked into blobs/ [doc].
  • An MLX folder has config.json with a quantization block (bits, group_size) or sits in an mlx-community repo; a plain model has config.json and *.safetensors without it. fleet models reports both, including folders outside the cache (~/models, LM Studio's folder).
  • Server side: GET /v1/models lists every cached repo as org/name plus the served path [doc][run].

Installed? Running?

  • Installed: fleet models lists mlx-lm under installed engines — on PATH or in ~/.grid/envs/mlx-lm. Install there and nowhere else, never into the system Python (a slow step: ask for the go-ahead first): uv venv --python 3.12 ~/.grid/envs/mlx-lm && uv pip install --python ~/.grid/envs/mlx-lm/bin/python mlx-lm [run] (pip install mlx-lm or conda install -c conda-forge mlx-lm [doc]). ENV below is that folder.
  • Running: fleet models → an engine on 8080 labelled openai-compatible (it has no owned_by) whose models are Hugging Face ids [run].

Start an already-downloaded model (never downloads)

Port P from outside machine.listeningPorts (8090 was taken by another app during testing [run]):

"$GRID_FLEET" serve mlx-P --env HF_HUB_OFFLINE=1 -- ENV/bin/mlx_lm.server \
  --model <snapshot dir or cached org/name> --host 127.0.0.1 --port P --max-tokens 32768 \
  --chat-template-args '{"enable_thinking":false}'
  • fleet serve starts it in its own session, so it outlives your shell; log in run/mlx-P.log, PID in run/mlx-P.pid. Under the agent's shell a plain nohup … & died with an empty log when the command returned, while a fleet serve process was still alive in a later command [run]. Never launchd.

  • HF_HUB_OFFLINE=1: no HTTP calls, cached files only, an error if missing [doc]. Without it an uncached model is downloaded from Hugging Face [doc].

  • --max-tokens defaults to 512 [doc]: a request without its own limit would stop mid-answer.

  • Sampling defaults are greedy (--temp 0.0, --top-p 1.0) [doc]; set --temp/--top-p/--top-k from the model card's recommended settings.

  • Ready log: Starting httpd at 127.0.0.1 on port P... [run].

  • There is no context size to set: the cache grows with the conversation, up to the model's max_position_embeddings. Check that it is ≥ 65536 and that weights + 64K of cache fit before starting.

Show full SKILL.md (374 more words)Show less

Ready means

Run "$GRID_FLEET" verify --at http://127.0.0.1:P/v1 --model <path or id> --kind mlx-lm right after serve — it waits for loading itself; no sleep, curl or log reading first. What it checks:

  1. The PID is alive. 2. GET /health → {"status": "ok"} [run] (in the source, not in SERVER.md [doc]).
  2. GET /v1/models → 200 [doc]. 4. One bounded /v1/chat/completions request, max_tokens 16, model = the path or id you started with → non-empty content [run].

Join Harness Compute

"$GRID_FLEET" run -- join GRID --at http://127.0.0.1:P/v1 -m <path or id you started with> --advertise-as ALIAS

/v1 is required: /models without it answers 404 (only chat accepts both) [doc][run]. Without --advertise-as the picker shows the whole snapshot path, and verify --alias waits five minutes for a name that never appears [run]. Grid's detector labels anything on 8080 as mlx, whatever it is [run].

Tool calls, JSON output, thinking

  • Tools: passed to the chat template; when the tokenizer has no tool-calling support the server logs "Received tools but model does not support tool calling" and answers without calls [doc]. Grid's probe reported no tool calls for a 0.5B 4-bit model whose template does mention tools [run]: run the tool check before offering a model for coding.
  • Thinking: --chat-template-args '{"enable_thinking":false}' at start, or chat_template_kwargs per request [doc]. Reasoning text comes back in message.reasoning [doc].
  • JSON: Grid's probe reported JSON object and schema modes as supported [run].

Memory and speed knobs

FlagDefaultUse
--decode-concurrency32requests decoded together [doc]
--prompt-concurrency8prompts prefilled together [doc]
--prefill-step-size2048lower it if prefill spikes memory [doc]
--prompt-cache-size / --prompt-cache-bytes10 caches / unlimitedcap memory kept for reuse [doc]
--kv-bits 4|8offsmaller cache for long context, but one request at a time [doc]; newer than 0.31.3 [run: absent from its help]
--draft-model, --num-draft-tokensnone, 3speculative decoding [doc]

Stop

"$GRID_FLEET" stop mlx-P, then confirm P is gone from fleet models --summary (ports in use).

Known failures → what to do

SignDo
OfflineModeIsEnabled or file-not-found at startthe model is not fully cached; offer the download as a slow step [doc]
answers stop around 512 tokensstart with --max-tokens [doc]
empty content, text in reasoningdisable thinking in the chat template [doc]
"Received tools but model does not support tool calling"not a coding model; pick another [doc]
address already in usepick another port; never stop the other process [run]

© autonomous-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in store/agents/autonomous-grid/skills/engine-mlx-lm of autonomous-ai/openharness.

Open the folder on GitHubat commit 50da5db

Compare with similar skills

Engine Mlx Lm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Engine Mlx Lm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Engine Mlx Lm this skillautonomous-ai/openharness1.1k—~1.9kAutomated safety check: PassMIT
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0
Huggingface LLM Trainerwaybarrios/opencode-power-pack533—~3kAutomated safety check: PassApache-2.0

Similar skills

  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check: notes
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Huggingface LLM Trainer

    waybarrios/opencode-power-pack

    Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.

    533 GitHub stars~3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from autonomous-ai/openharness

All 100 skills in this repo
  • G-code Slicer Tool

    autonomous-ai/openharness

    Slices 3D mesh files into printer-profiled plain G-code through real slicer CLIs, with backend discovery, input inspection, dry runs and static validation.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Home Assistant Automation Builder

    autonomous-ai/openharness

    Turns a home-automation request into standard, testable automations.yaml, run against Home Assistant Core's real triggers and verified with its own trace tool.

    1.1k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Score Music Composer

    autonomous-ai/openharness

    Turns a musical brief into LilyPond concert-pitch music, checked parts for each instrument and a playable practice pack.

    1.1k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • OrcaSlicer 3MF and G-code Workflow

    autonomous-ai/openharness

    Turns an STL and explicit printer and material requirements into compared OrcaSlicer plans, an editable 3MF project, checked G-code and a portable handoff.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Sheets and Docs Report Builder

    autonomous-ai/openharness

    Builds an editable DOCX report, a formula-driven XLSX workbook and a fresh LibreOffice PDF preview from one structured source file, then checks them together.

    1.1k GitHub stars~708 tokensUpdated today
    Auto-check passed
  • Bambu Labs

    autonomous-ai/openharness

    Dry-run, upload, and cautiously initiate local Bambu Lab print jobs from validated plain .gcode, using Bambu LAN FTPS/MQTT handoffs.

    1.1k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check: warnings

Questions about Engine Mlx Lm

What does Engine Mlx Lm do?

Serve an MLX or Hugging Face safetensors model already on this Mac with mlx-lm's server and join it to the person's fleet. Engine Mlx Lm is an agent skill from autonomous-ai/openharness. Serve an MLX or Hugging Face safetensors model already on this Mac with mlx-lm's server and join it to the person's fleet.

When should I use Engine Mlx Lm?

Engine Mlx Lm fits situations like: tasks that involve Model hubs and datasets.

How do I install Engine Mlx Lm in Claude Code?

Run `npx skills add autonomous-ai/openharness --skill engine-mlx-lm -a claude-code`. Or copy the skill folder (store/agents/autonomous-grid/skills/engine-mlx-lm in autonomous-ai/openharness) into .claude/skills/engine-mlx-lm in your project. Claude Code loads it when a task matches its description.

How do I install Engine Mlx Lm in Codex?

Run `npx skills add autonomous-ai/openharness --skill engine-mlx-lm -a codex`. Or copy the skill folder (store/agents/autonomous-grid/skills/engine-mlx-lm in autonomous-ai/openharness) into .agents/skills/engine-mlx-lm in your project. Codex loads it when a task matches its description.

Can I use Engine Mlx Lm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/openharness --skill engine-mlx-lm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/engine-mlx-lm, .gemini/skills/engine-mlx-lm, .github/skills/engine-mlx-lm and .opencode/skills/engine-mlx-lm in your project.

What does Engine Mlx Lm need to run?

Going by SKILL.md and its folder, Engine Mlx Lm needs the command-line tools its instructions call (uv, pip and conda). Our summary lists: Python 3.

Does Engine Mlx Lm access the network?

SKILL.md names 2 domains. As links in the text: github.com and huggingface.co. This is read from the text; nothing was executed.

Is Engine Mlx Lm safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Engine Mlx Lm use?

Engine Mlx Lm is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Engine Mlx Lm use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Engine Mlx Lm?

Skills that share tags, products or a category with Engine Mlx Lm: Hugging Face LLM Trainer (huggingface/skills, 11k stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Add Model (guoqingbao/xinfer, 333 stars) and Hugging Face Local Models (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Engine Mlx Lm?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/openharness, which has 1,149 GitHub stars. The repository holds 100 skills in this directory. The repository was last updated on October 8, 2026.

Source: autonomous-ai/openharness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.