Agent skill

LLM Integration

by yonatangross in yonatangross/orchestkit

LLM integration patterns for function calling, streaming responses, local inference with Ollama, and fine-tuning customization.

MITAuto-check passedAI & LLM Engineering

Install LLM Integration

skills CLI
$ npx skills add yonatangross/orchestkit --skill llm-integration -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yonatangross/orchestkit llm-integration --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/llm-integration .claude/skills/llm-integration && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-integration
GitHub stars
290
Token cost
~2.7k tokens
SKILL.md length
900 words
Files
36 (incl. scripts, references)
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

LLM integration patterns for function calling, streaming responses, local inference with Ollama, and fine-tuning customization.

  • Implementing tool use
  • SKILL.md covers Quick Reference, Quick Start, Function Calling and Streaming, plus 10 more sections
  • Local model deployment
  • LoRA/QLoRA fine-tuning

What it does

LLM Integration is an agent skill from yonatangross/orchestkit. LLM integration patterns for function calling, streaming responses, local inference with Ollama, and fine-tuning customization. Use when implementing tool use, SSE streaming, local model deployment, LoRA/QLoRA fine-tuning, or multi-provider LLM APIs.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 38 other files, including scripts and reference files (for example `checklists/fine-tuning-decision.md`, `checklists/streaming-checklist.md` and `checklists/tool-checklist.md`). Compatibility notes: Claude Code 2.1.277+.

It sits in AI & LLM Engineering, covering Fine-tuning, Structured output and tool calling and LLM API integration. It works with Ollama. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.

When your agent uses it

  • Implementing tool use
  • Local model deployment
  • LoRA/QLoRA fine-tuning
  • Multi-provider LLM APIs

Example prompts

  • “/llm-integration”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Claude Code 2.1.277+.
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, WebFetch, WebSearch

What it can do on your machine

Read from SKILL.md and the folder at commit 02bbf9a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • WebFetch
    • WebSearch

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.openai.com
    • huggingface.co
    • platform.claude.com
    • docs.claude.com
    • developer.mozilla.org
    • docs.unsloth.ai
    • sbert.net

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Claude Code 2.1.277+.

    From compatibility in the SKILL.md frontmatter.

Context cost

LLM Integration loads about 2.7k tokens when it runs, and up to ~4.3k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 900 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from yonatangross/orchestkit at commit 02bbf9a, republished under its MIT licence (© yonatangross). 900 words, ~2,717 tokens.

Download SKILL.mdSave it as .claude/skills/llm-integration/SKILL.md (or your agent's skills folder). This skill also uses 35 other files; get the full folder from GitHub.
name
llm-integration
description
LLM integration patterns for function calling, streaming responses, local inference with Ollama, and fine-tuning customization. Use when implementing tool use, SSE streaming, local model deployment, LoRA/QLoRA fine-tuning, or multi-provider LLM APIs.
allowed-tools
Read, Glob, Grep, WebFetch, WebSearch
compatibility
Claude Code 2.1.277+.
license
MIT
user-invocable
false
disable-model-invocation
true
metadata.owner-agent
llm-integrator
metadata.category
mcp-enhancement
metadata.version
2.0.0
metadata.author
OrchestKit
metadata.complexity
medium
metadata.tags
llm, function-calling, streaming, ollama, fine-tuning, lora, tool-use, local-inference

LLM Integration

Patterns for integrating LLMs into production applications: tool use, streaming, local inference, and fine-tuning. Each category has individual rule files in rules/ loaded on-demand.

Quick Reference

CategoryRulesImpactWhen to Use
Function Calling3CRITICALTool definitions, parallel execution, input validation
Streaming3HIGHSSE endpoints, structured streaming, backpressure handling
Local Inference3HIGHOllama setup, model selection, GPU optimization
Fine-Tuning3HIGHLoRA/QLoRA training, dataset preparation, evaluation
Context Optimization2HIGHWindow management, compression, caching, budget scaling
Evaluation2HIGHLLM-as-judge, RAGAS metrics, quality gates, benchmarks
Prompt Engineering4HIGHCoT, few-shot, versioning, DSPy optimization, ReAct, cost optimization

Total: 20 rules across 7 categories

Quick Start

python
# Function calling: strict mode tool definition
tools = [{
    "type": "function",
    "function": {
        "name": "search_documents",
        "description": "Search knowledge base",
        "strict": True,
        "parameters": {
            "type": "object",
            "properties": {
                "query": {"type": "string", "description": "Search query"},
                "limit": {"type": "integer", "description": "Max results"}
            },
            "required": ["query", "limit"],
            "additionalProperties": False
        }
    }
}]
python
# Streaming: SSE endpoint with FastAPI
@app.get("/chat/stream")
async def stream_chat(prompt: str):
    async def generate():
        async for token in async_stream(prompt):
            yield {"event": "token", "data": token}
        yield {"event": "done", "data": ""}
    return EventSourceResponse(generate())
python
# Local inference: Ollama with LangChain
llm = ChatOllama(
    model="deepseek-r1:70b",
    base_url="http://localhost:11434",
    temperature=0.0,
    num_ctx=32768,
)
python
# Fine-tuning: QLoRA with Unsloth
model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/Meta-Llama-3.1-8B",
    max_seq_length=2048, load_in_4bit=True,
)
model = FastLanguageModel.get_peft_model(model, r=16, lora_alpha=32)

Function Calling

Enable LLMs to use external tools and return structured data. Use strict mode schemas (2026 best practice) for reliability. On Claude keep the tools array stable across turns so the prompt cache holds; add tools mid-session with an inline tool_addition (beta inline-tools-2026-09-15, Claude API only, available on Sonnet 5.5 but not on Sonnet 5) instead of swapping subsets (https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#mid-conversation-tool-changes). For large tool sets, use a non-deferred tool search tool with defer_loading: true on the rest, never defer_loading alone (a request cannot have every tool deferred). Forced tool_choice (any or a named tool) returns a 400 on Sonnet 5.5, Opus 5.5, Fable 5.1 and Mythos 5.1; use auto with strict: true tools. Validate all inputs with Pydantic/Zod, and return errors as tool results. <!-- model-recency-ok: per-model contrast, Sonnet 5 lacks what Sonnet 5.5 has (platform docs 2026-09-28) -->

  • calling-tool-definition.md -- Strict mode schemas, OpenAI/Anthropic formats, LangChain binding
  • calling-parallel.md -- Parallel tool execution, asyncio.gather, strict mode constraints
  • calling-validation.md -- Input validation, error handling, tool execution loops

Streaming

Deliver LLM responses in real-time for better UX. Use SSE for web, WebSocket for bidirectional. Handle backpressure with bounded queues.

  • streaming-sse.md -- FastAPI SSE endpoints, frontend consumers, async iterators
  • streaming-structured.md -- Streaming with tool calls, partial JSON parsing, chunk accumulation
  • streaming-backpressure.md -- Backpressure handling, bounded buffers, cancellation

Local Inference

Run LLMs locally with Ollama for cost savings (93% vs cloud), privacy, and offline development. Pre-warm models, use provider factory for cloud/local switching.

  • local-ollama-setup.md -- Installation, model pulling, environment configuration
  • local-model-selection.md -- Model comparison by task, hardware profiles, quantization
  • local-gpu-optimization.md -- Apple Silicon tuning, keep-alive, CI integration

Fine-Tuning

Customize LLMs with parameter-efficient techniques. Fine-tune ONLY after exhausting prompt engineering and RAG. Requires 1000+ quality examples.

  • tuning-lora.md -- LoRA/QLoRA configuration, Unsloth training, adapter merging
  • tuning-dataset-prep.md -- Synthetic data generation, quality validation, deduplication
  • tuning-evaluation.md -- DPO alignment, evaluation metrics, anti-patterns

Context Optimization

Manage context windows, compression, and attention-aware positioning. Optimize for tokens-per-task.

  • context-window-management.md -- Five-layer architecture, anchored summarization, compression triggers
  • context-caching.md -- Just-in-time loading, budget scaling, probe evaluation, CC 2.1.32+

Evaluation

Evaluate LLM outputs with multi-dimension scoring, quality gates, and benchmarks.

  • evaluation-metrics.md -- LLM-as-judge, RAGAS metrics, hallucination detection
  • evaluation-benchmarks.md -- Quality gates, batch evaluation, pairwise comparison

Prompt Engineering

Design, version, and optimize prompts for production LLM applications.

  • prompt-design.md -- Chain-of-Thought, few-shot learning, pattern selection guide
  • prompt-testing.md -- Langfuse versioning, DSPy optimization, A/B testing, self-consistency
  • prompt-react-pattern.md -- ReAct loop for tool-using agents, thought-action-observation format
  • prompt-optimization.md -- Token reduction, cost optimization, model tiering, prompt spec format
Show full SKILL.md (417 more words)Show less

Upstream coverage (do not restate)

These topics are covered by their vendors' own documentation. This skill points at them instead of teaching them; the rules above keep only our floors, ceilings and scars. Our delta on all of it is in references/ork-delta.md.

TopicFirst-party source
Strict-mode tool schemas, structured outputshttps://platform.openai.com/docs/guides/function-calling
Anthropic input_schema / tool_usehttps://docs.claude.com/en/docs/agents-and-tools/tool-use/overview
SSE client mechanics, reconnection, cancellationhttps://developer.mozilla.org/en-US/docs/Web/API/EventSource
LoRA / QLoRA config, target modules, adapter merginghttps://huggingface.co/docs/peft/developer_guides/lora
Unsloth training loop, 4-bit loadinghttps://docs.unsloth.ai/get-started/fine-tuning-llms-guide
DPO, preference pairs, beta tuning, RLHF comparisonhttps://huggingface.co/docs/trl/dpo_trainer
SFT dataset formats (Alpaca, ChatML)https://huggingface.co/docs/trl/sft_trainer
Embedding similarity for dataset deduphttps://sbert.net/
Fine-tune vs prompt vs RAG decision frameworkhttps://platform.openai.com/docs/guides/optimizing-llm-accuracy
Vendor token pricing (never hardcode it here)https://platform.openai.com/docs/pricing

Supporting Files

  • references/ork-delta.md -- our delta: scars, house ceilings, retired-file provenance
  • references/model-selection.md -- local model comparison by task and hardware
  • scripts/create-lora-config.md -- LoRA config scaffold with auto-detected model type

Key Decisions

DecisionRecommendation
Tool schema modestrict: true (2026 best practice)
Tool countLoad roughly 5-15 up front; defer the rest behind a non-deferred tool search tool
Streaming protocolSSE for web, WebSocket for bidirectional
Buffer size50-200 tokens
Local model (reasoning)deepseek-r1:70b
Local model (coding)qwen2.5-coder:32b
Fine-tuning approachLoRA/QLoRA (try prompting first)
LoRA rank16-64 typical
Training epochs1-3 (more risks overfitting)
Context compressionAnchored iterative (60-80%)
Compress trigger70% utilization, target 50%
Judge modelclaude-haiku-4-5-20251001 (cost tier), gpt-5.5, or gemini-3.8-flash (Google's GA general-purpose model, priced under Haiku through 2026-12-31 and a third vendor for the different-model rule; price lives in models.vocab.json, not here)
Quality threshold0.7 production, 0.6 drafts
Few-shot examples3-5 diverse, representative
Prompt versioningLangfuse with labels
Auto-optimizationDSPy MIPROv2
  • ork:rag-retrieval -- Embedding patterns, when RAG is better than fine-tuning
  • agent-loops -- Multi-step tool use with reasoning
  • llm-evaluation -- Evaluate fine-tuned and local models
  • langfuse-observability -- Track training experiments

Capability Details

function-calling

Keywords: tool, function, define tool, tool schema, function schema, strict mode, parallel tools Solves:

  • Define tools with clear descriptions and strict schemas
  • Execute tool calls in parallel with asyncio.gather
  • Validate inputs and handle errors in tool execution loops
streaming

Keywords: streaming, SSE, Server-Sent Events, real-time, backpressure, token stream Solves:

  • Stream LLM tokens via SSE endpoints
  • Handle tool calls within streams
  • Manage backpressure with bounded queues
local-inference

Keywords: Ollama, local, self-hosted, model selection, GPU, Apple Silicon Solves:

  • Set up Ollama for local LLM inference
  • Select models based on task and hardware
  • Optimize GPU usage and CI integration
fine-tuning

Keywords: LoRA, QLoRA, fine-tune, DPO, synthetic data, PEFT, alignment Solves:

  • Configure LoRA/QLoRA for parameter-efficient training
  • Generate and validate synthetic training data
  • Align models with DPO and evaluate results

© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 35 other files (scripts, references) in src/skills/llm-integration of yonatangross/orchestkit.

  • SKILL.md
  • checklists/fine-tuning-decision.md
  • checklists/streaming-checklist.md
  • checklists/tool-checklist.md
  • metadata.json
  • references/model-selection.md
  • references/ork-delta.md
  • rules/_sections.md
  • rules/_template.md
  • rules/calling-parallel.md
  • rules/calling-tool-definition.md
  • rules/calling-validation.md
  • rules/context-caching.md
  • rules/context-window-management.md
  • rules/evaluation-benchmarks.md
  • rules/evaluation-metrics.md
  • rules/local-gpu-optimization.md
  • rules/local-model-selection.md
  • … and 18 more

Open the folder on GitHubat commit 02bbf9a

Compare with similar skills

LLM Integration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Integration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Integration this skillyonatangross/orchestkit290—~2.7kAutomated safety check: PassMIT
Swift Mlx Lmkellyvv/PhoneClaw1.3k—~3.7kAutomated safety check: PassApache-2.0
Tanstack AIsecondsky/claude-skills227—~3.6kAutomated safety check: NotesMIT
Hugging Face LLM Trainerhuggingface/skills11k1 repos~7.2kAutomated safety check: PassApache-2.0
Agnes Free Textkangarooking/agnes-free-model-skills199—~630Automated safety check: PassMIT
Instructor Structured LLM OutputsOrchestra-Research/AI-Research-SKILLs13k6 repos~4.2kAutomated safety check: PassMIT

Similar skills

  • Swift Mlx Lm

    kellyvv/PhoneClaw

    MLX Swift LM - Run LLMs and VLMs on Apple Silicon using MLX.

    1.3k GitHub stars~3.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Tanstack AI

    secondsky/claude-skills

    TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama.

    227 GitHub stars~3.6k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check: notes
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 1 repo~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Agnes Free Text

    kangarooking/agnes-free-model-skills

    Call the free Agnes text model API for chat completions, streaming answers, coding help, tool-calling experiments, and OpenAI-compatible text generation.

    199 GitHub stars~630 tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Instructor Structured LLM Outputs

    Orchestra-Research/AI-Research-SKILLs

    Shows how to pull validated, typed data out of LLM responses with Instructor and Pydantic models, including retries on failure and partial streaming.

    13k GitHub starsUsed in 6 repos~4.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Trl Post Training

    burtenshaw/training-agents

    A skill your agent uses when building, reviewing, or editing TRL post-training workflows for agentic applications, including SFT, DPO, GRPO, RLOO, reward modeling, dataset formats, chat templates…

    153 GitHub stars~599 tokensUpdated 26 days ago
    AI & LLM EngineeringAuto-check passed

More from yonatangross/orchestkit

All 108 skills in this repo
  • API Design

    yonatangross/orchestkit

    API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.

    290 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Architecture Decision Record

    yonatangross/orchestkit

    ADR templates in the Nygard format with context, decision, consequences, and alternatives.

    290 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Audit Full

    yonatangross/orchestkit

    Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.

    290 GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Code Review Playbook

    yonatangross/orchestkit

    Structured review processes, conventional comments, language-specific checklists, and feedback templates.

    290 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Create PR

    yonatangross/orchestkit

    Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.

    290 GitHub stars~4.5k tokensUpdated today
    Auto-check: notes
  • Explore

    yonatangross/orchestkit

    Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.

    290 GitHub stars~3.9k tokensUpdated today
    Auto-check: notes

Works with

Questions about LLM Integration

What does LLM Integration do?

LLM integration patterns for function calling, streaming responses, local inference with Ollama, and fine-tuning customization. LLM Integration is an agent skill from yonatangross/orchestkit. LLM integration patterns for function calling, streaming responses, local inference with Ollama, and fine-tuning customization.

When should I use LLM Integration?

LLM Integration fits situations like: implementing tool use; local model deployment; loRA/QLoRA fine-tuning; multi-provider LLM APIs.

How do I install LLM Integration in Claude Code?

Run `npx skills add yonatangross/orchestkit --skill llm-integration -a claude-code`. Or copy the skill folder (src/skills/llm-integration in yonatangross/orchestkit) into .claude/skills/llm-integration in your project. Claude Code loads it when a task matches its description.

How do I install LLM Integration in Codex?

Run `npx skills add yonatangross/orchestkit --skill llm-integration -a codex`. Or copy the skill folder (src/skills/llm-integration in yonatangross/orchestkit) into .agents/skills/llm-integration in your project. Codex loads it when a task matches its description.

Can I use LLM Integration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill llm-integration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-integration, .gemini/skills/llm-integration, .github/skills/llm-integration and .opencode/skills/llm-integration in your project.

What does LLM Integration need to run?

SKILL.md names no scripts, command-line tools or credentials: LLM Integration is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Glob, Grep, WebFetch, WebSearch. Compatibility (from SKILL.md): Claude Code 2.1.277+..

Does LLM Integration access the network?

SKILL.md names 7 domains. As links in the text: platform.openai.com, huggingface.co, platform.claude.com, docs.claude.com, developer.mozilla.org, docs.unsloth.ai and sbert.net. This is read from the text; nothing was executed.

Is LLM Integration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does LLM Integration use?

LLM Integration is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Integration use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to LLM Integration?

Skills that share tags, products or a category with LLM Integration: Swift Mlx Lm (kellyvv/PhoneClaw, 1.3k stars), Tanstack AI (secondsky/claude-skills, 227 stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and Agnes Free Text (kangarooking/agnes-free-model-skills, 199 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Integration?

yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 290 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.

Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.