Swift Mlx Lm
kellyvv/PhoneClaw
MLX Swift LM - Run LLMs and VLMs on Apple Silicon using MLX.
LLM integration patterns for function calling, streaming responses, local inference with Ollama, and fine-tuning customization.
$ npx skills add yonatangross/orchestkit --skill llm-integration -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install yonatangross/orchestkit llm-integration --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/llm-integration .claude/skills/llm-integration && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "llm-integration" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/llm-integration into .claude/skills/llm-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-integration", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/yonatangross/orchestkit/tree/main/src/skills/llm-integrationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add yonatangross/orchestkit --skill llm-integration -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install yonatangross/orchestkit llm-integration --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/skills/llm-integration .agents/skills/llm-integration && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "llm-integration" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/llm-integration into .agents/skills/llm-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-integration", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill llm-integration -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install yonatangross/orchestkit llm-integration --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/skills/llm-integration .cursor/skills/llm-integration && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "llm-integration" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/llm-integration into .cursor/skills/llm-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-integration", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/yonatangross/orchestkit.git --path src/skills/llm-integration--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add yonatangross/orchestkit --skill llm-integration -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install yonatangross/orchestkit llm-integration --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/skills/llm-integration .gemini/skills/llm-integration && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "llm-integration" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/llm-integration into .gemini/skills/llm-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-integration", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install yonatangross/orchestkit llm-integrationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add yonatangross/orchestkit --skill llm-integration -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/skills/llm-integration .github/skills/llm-integration && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "llm-integration" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/llm-integration into .github/skills/llm-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-integration", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill llm-integration -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install yonatangross/orchestkit llm-integration --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/skills/llm-integration .opencode/skills/llm-integration && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "llm-integration" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/llm-integration into .opencode/skills/llm-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-integration", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
llm-integrationLLM integration patterns for function calling, streaming responses, local inference with Ollama, and fine-tuning customization.
LLM Integration is an agent skill from yonatangross/orchestkit. LLM integration patterns for function calling, streaming responses, local inference with Ollama, and fine-tuning customization. Use when implementing tool use, SSE streaming, local model deployment, LoRA/QLoRA fine-tuning, or multi-provider LLM APIs.
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 38 other files, including scripts and reference files (for example `checklists/fine-tuning-decision.md`, `checklists/streaming-checklist.md` and `checklists/tool-checklist.md`). Compatibility notes: Claude Code 2.1.277+.
It sits in AI & LLM Engineering, covering Fine-tuning, Structured output and tool calling and LLM API integration. It works with Ollama. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.
Read from SKILL.md and the folder at commit 02bbf9a. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGlobGrepWebFetchWebSearchFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
platform.openai.comhuggingface.coplatform.claude.comdocs.claude.comdeveloper.mozilla.orgdocs.unsloth.aisbert.netFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Claude Code 2.1.277+.
From compatibility in the SKILL.md frontmatter.
LLM Integration loads about 2.7k tokens when it runs, and up to ~4.3k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 900 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from yonatangross/orchestkit at commit 02bbf9a, republished under its MIT licence (© yonatangross). 900 words, ~2,717 tokens.
.claude/skills/llm-integration/SKILL.md (or your agent's skills folder). This skill also uses 35 other files; get the full folder from GitHub.Patterns for integrating LLMs into production applications: tool use, streaming, local inference, and fine-tuning. Each category has individual rule files in rules/ loaded on-demand.
| Category | Rules | Impact | When to Use |
|---|---|---|---|
| Function Calling | 3 | CRITICAL | Tool definitions, parallel execution, input validation |
| Streaming | 3 | HIGH | SSE endpoints, structured streaming, backpressure handling |
| Local Inference | 3 | HIGH | Ollama setup, model selection, GPU optimization |
| Fine-Tuning | 3 | HIGH | LoRA/QLoRA training, dataset preparation, evaluation |
| Context Optimization | 2 | HIGH | Window management, compression, caching, budget scaling |
| Evaluation | 2 | HIGH | LLM-as-judge, RAGAS metrics, quality gates, benchmarks |
| Prompt Engineering | 4 | HIGH | CoT, few-shot, versioning, DSPy optimization, ReAct, cost optimization |
Total: 20 rules across 7 categories
# Function calling: strict mode tool definition
tools = [{
"type": "function",
"function": {
"name": "search_documents",
"description": "Search knowledge base",
"strict": True,
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "Search query"},
"limit": {"type": "integer", "description": "Max results"}
},
"required": ["query", "limit"],
"additionalProperties": False
}
}
}]# Streaming: SSE endpoint with FastAPI
@app.get("/chat/stream")
async def stream_chat(prompt: str):
async def generate():
async for token in async_stream(prompt):
yield {"event": "token", "data": token}
yield {"event": "done", "data": ""}
return EventSourceResponse(generate())# Local inference: Ollama with LangChain
llm = ChatOllama(
model="deepseek-r1:70b",
base_url="http://localhost:11434",
temperature=0.0,
num_ctx=32768,
)# Fine-tuning: QLoRA with Unsloth
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="unsloth/Meta-Llama-3.1-8B",
max_seq_length=2048, load_in_4bit=True,
)
model = FastLanguageModel.get_peft_model(model, r=16, lora_alpha=32)Enable LLMs to use external tools and return structured data. Use strict mode schemas (2026 best practice) for reliability. On Claude keep the tools array stable across turns so the prompt cache holds; add tools mid-session with an inline tool_addition (beta inline-tools-2026-09-15, Claude API only, available on Sonnet 5.5 but not on Sonnet 5) instead of swapping subsets (https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#mid-conversation-tool-changes). For large tool sets, use a non-deferred tool search tool with defer_loading: true on the rest, never defer_loading alone (a request cannot have every tool deferred). Forced tool_choice (any or a named tool) returns a 400 on Sonnet 5.5, Opus 5.5, Fable 5.1 and Mythos 5.1; use auto with strict: true tools. Validate all inputs with Pydantic/Zod, and return errors as tool results. <!-- model-recency-ok: per-model contrast, Sonnet 5 lacks what Sonnet 5.5 has (platform docs 2026-09-28) -->
calling-tool-definition.md -- Strict mode schemas, OpenAI/Anthropic formats, LangChain bindingcalling-parallel.md -- Parallel tool execution, asyncio.gather, strict mode constraintscalling-validation.md -- Input validation, error handling, tool execution loopsDeliver LLM responses in real-time for better UX. Use SSE for web, WebSocket for bidirectional. Handle backpressure with bounded queues.
streaming-sse.md -- FastAPI SSE endpoints, frontend consumers, async iteratorsstreaming-structured.md -- Streaming with tool calls, partial JSON parsing, chunk accumulationstreaming-backpressure.md -- Backpressure handling, bounded buffers, cancellationRun LLMs locally with Ollama for cost savings (93% vs cloud), privacy, and offline development. Pre-warm models, use provider factory for cloud/local switching.
local-ollama-setup.md -- Installation, model pulling, environment configurationlocal-model-selection.md -- Model comparison by task, hardware profiles, quantizationlocal-gpu-optimization.md -- Apple Silicon tuning, keep-alive, CI integrationCustomize LLMs with parameter-efficient techniques. Fine-tune ONLY after exhausting prompt engineering and RAG. Requires 1000+ quality examples.
tuning-lora.md -- LoRA/QLoRA configuration, Unsloth training, adapter mergingtuning-dataset-prep.md -- Synthetic data generation, quality validation, deduplicationtuning-evaluation.md -- DPO alignment, evaluation metrics, anti-patternsManage context windows, compression, and attention-aware positioning. Optimize for tokens-per-task.
context-window-management.md -- Five-layer architecture, anchored summarization, compression triggerscontext-caching.md -- Just-in-time loading, budget scaling, probe evaluation, CC 2.1.32+Evaluate LLM outputs with multi-dimension scoring, quality gates, and benchmarks.
evaluation-metrics.md -- LLM-as-judge, RAGAS metrics, hallucination detectionevaluation-benchmarks.md -- Quality gates, batch evaluation, pairwise comparisonDesign, version, and optimize prompts for production LLM applications.
prompt-design.md -- Chain-of-Thought, few-shot learning, pattern selection guideprompt-testing.md -- Langfuse versioning, DSPy optimization, A/B testing, self-consistencyprompt-react-pattern.md -- ReAct loop for tool-using agents, thought-action-observation formatprompt-optimization.md -- Token reduction, cost optimization, model tiering, prompt spec formatThese topics are covered by their vendors' own documentation. This skill points
at them instead of teaching them; the rules above keep only our floors, ceilings
and scars. Our delta on all of it is in references/ork-delta.md.
| Topic | First-party source |
|---|---|
| Strict-mode tool schemas, structured outputs | https://platform.openai.com/docs/guides/function-calling |
Anthropic input_schema / tool_use | https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview |
| SSE client mechanics, reconnection, cancellation | https://developer.mozilla.org/en-US/docs/Web/API/EventSource |
| LoRA / QLoRA config, target modules, adapter merging | https://huggingface.co/docs/peft/developer_guides/lora |
| Unsloth training loop, 4-bit loading | https://docs.unsloth.ai/get-started/fine-tuning-llms-guide |
| DPO, preference pairs, beta tuning, RLHF comparison | https://huggingface.co/docs/trl/dpo_trainer |
| SFT dataset formats (Alpaca, ChatML) | https://huggingface.co/docs/trl/sft_trainer |
| Embedding similarity for dataset dedup | https://sbert.net/ |
| Fine-tune vs prompt vs RAG decision framework | https://platform.openai.com/docs/guides/optimizing-llm-accuracy |
| Vendor token pricing (never hardcode it here) | https://platform.openai.com/docs/pricing |
references/ork-delta.md -- our delta: scars, house ceilings, retired-file provenancereferences/model-selection.md -- local model comparison by task and hardwarescripts/create-lora-config.md -- LoRA config scaffold with auto-detected model type| Decision | Recommendation |
|---|---|
| Tool schema mode | strict: true (2026 best practice) |
| Tool count | Load roughly 5-15 up front; defer the rest behind a non-deferred tool search tool |
| Streaming protocol | SSE for web, WebSocket for bidirectional |
| Buffer size | 50-200 tokens |
| Local model (reasoning) | deepseek-r1:70b |
| Local model (coding) | qwen2.5-coder:32b |
| Fine-tuning approach | LoRA/QLoRA (try prompting first) |
| LoRA rank | 16-64 typical |
| Training epochs | 1-3 (more risks overfitting) |
| Context compression | Anchored iterative (60-80%) |
| Compress trigger | 70% utilization, target 50% |
| Judge model | claude-haiku-4-5-20251001 (cost tier), gpt-5.5, or gemini-3.8-flash (Google's GA general-purpose model, priced under Haiku through 2026-12-31 and a third vendor for the different-model rule; price lives in models.vocab.json, not here) |
| Quality threshold | 0.7 production, 0.6 drafts |
| Few-shot examples | 3-5 diverse, representative |
| Prompt versioning | Langfuse with labels |
| Auto-optimization | DSPy MIPROv2 |
ork:rag-retrieval -- Embedding patterns, when RAG is better than fine-tuningagent-loops -- Multi-step tool use with reasoningllm-evaluation -- Evaluate fine-tuned and local modelslangfuse-observability -- Track training experimentsKeywords: tool, function, define tool, tool schema, function schema, strict mode, parallel tools Solves:
Keywords: streaming, SSE, Server-Sent Events, real-time, backpressure, token stream Solves:
Keywords: Ollama, local, self-hosted, model selection, GPU, Apple Silicon Solves:
Keywords: LoRA, QLoRA, fine-tune, DPO, synthetic data, PEFT, alignment Solves:
© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 35 other files (scripts, references) in src/skills/llm-integration of yonatangross/orchestkit.
Open the folder on GitHubat commit 02bbf9a
LLM Integration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| LLM Integration this skillyonatangross/orchestkit | 290 | — | ~2.7k | Automated safety check: Pass | MIT | |
| Swift Mlx Lmkellyvv/PhoneClaw | 1.3k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | |
| Tanstack AIsecondsky/claude-skills | 227 | — | ~3.6k | Automated safety check: Notes | MIT | |
| Hugging Face LLM Trainerhuggingface/skills | 11k | 1 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| Agnes Free Textkangarooking/agnes-free-model-skills | 199 | — | ~630 | Automated safety check: Pass | MIT | |
| Instructor Structured LLM OutputsOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~4.2k | Automated safety check: Pass | MIT |
kellyvv/PhoneClaw
MLX Swift LM - Run LLMs and VLMs on Apple Silicon using MLX.
secondsky/claude-skills
TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama.
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
kangarooking/agnes-free-model-skills
Call the free Agnes text model API for chat completions, streaming answers, coding help, tool-calling experiments, and OpenAI-compatible text generation.
Orchestra-Research/AI-Research-SKILLs
Shows how to pull validated, typed data out of LLM responses with Instructor and Pydantic models, including retries on failure and partial streaming.
burtenshaw/training-agents
A skill your agent uses when building, reviewing, or editing TRL post-training workflows for agentic applications, including SFT, DPO, GRPO, RLOO, reward modeling, dataset formats, chat templates…
yonatangross/orchestkit
API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.
yonatangross/orchestkit
ADR templates in the Nygard format with context, decision, consequences, and alternatives.
yonatangross/orchestkit
Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.
yonatangross/orchestkit
Structured review processes, conventional comments, language-specific checklists, and feedback templates.
yonatangross/orchestkit
Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.
yonatangross/orchestkit
Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.
Works with
Categories
LLM integration patterns for function calling, streaming responses, local inference with Ollama, and fine-tuning customization. LLM Integration is an agent skill from yonatangross/orchestkit. LLM integration patterns for function calling, streaming responses, local inference with Ollama, and fine-tuning customization.
LLM Integration fits situations like: implementing tool use; local model deployment; loRA/QLoRA fine-tuning; multi-provider LLM APIs.
Run `npx skills add yonatangross/orchestkit --skill llm-integration -a claude-code`. Or copy the skill folder (src/skills/llm-integration in yonatangross/orchestkit) into .claude/skills/llm-integration in your project. Claude Code loads it when a task matches its description.
Run `npx skills add yonatangross/orchestkit --skill llm-integration -a codex`. Or copy the skill folder (src/skills/llm-integration in yonatangross/orchestkit) into .agents/skills/llm-integration in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill llm-integration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-integration, .gemini/skills/llm-integration, .github/skills/llm-integration and .opencode/skills/llm-integration in your project.
SKILL.md names no scripts, command-line tools or credentials: LLM Integration is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Glob, Grep, WebFetch, WebSearch. Compatibility (from SKILL.md): Claude Code 2.1.277+..
SKILL.md names 7 domains. As links in the text: platform.openai.com, huggingface.co, platform.claude.com, docs.claude.com, developer.mozilla.org, docs.unsloth.ai and sbert.net. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
LLM Integration is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with LLM Integration: Swift Mlx Lm (kellyvv/PhoneClaw, 1.3k stars), Tanstack AI (secondsky/claude-skills, 227 stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and Agnes Free Text (kangarooking/agnes-free-model-skills, 199 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 290 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.
Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.