Agent skill

Langchain Model Inference

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Invoke Claude, GPT-4o, and Gemini through LangChain 1.0 without tripping on the content-block, token-accounting, and structured-output quirks that silently break production code.

MITAuto-check passedAI & LLM Engineering

Install Langchain Model Inference

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-model-inference -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-model-inference --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/langchain-model-inference .claude/skills/langchain-model-inference && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langchain-model-inference
GitHub stars
2.8k
Token cost
~2.7k tokens
SKILL.md length
784 words
Files
6 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Invoke Claude, GPT-4o, and Gemini through LangChain 1.0 without tripping on the content-block, token-accounting, and structured-output quirks that silently break production code.

  • Works in 5 steps: Initialize a chat model with explicit,… → Iterate AIMessage.content as typed… → Route across providers with a factory,… → …
  • Initializing chat models
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
  • Calls pip; needs ANTHROPIC_API_KEY and OPENAI_API_KEY

What it does

Langchain Model Inference is an agent skill from jeremylongshore/tons-of-skills-marketplace. Invoke Claude, GPT-4o, and Gemini through LangChain 1.0 without tripping on the content-block, token-accounting, and structured-output quirks that silently break production code. Use when initializing chat models, routing across providers, iterating AIMessage content, or choosing a structured-output method. Trigger with "langchain model inference", "ChatAnthropic", "ChatOpenAI", "withstructuredoutput", "AIMessage content blocks", "langchain routing".

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/content-blocks.md`, `references/one-pager.md` and `references/provider-quirks.md`). Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering Building AI agents and Structured output and tool calling. It works with LangChain and OpenAI. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Initializing chat models
  • Routing across providers
  • Iterating AIMessage content
  • Choosing a structured-output method

Example prompts

  • “langchain model inference”
  • “ChatAnthropic”
  • “ChatOpenAI”
  • “/langchain-model-inference”

Requirements

  • Python 3
  • A credential in ANTHROPIC_API_KEY
  • A credential in OPENAI_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(python:*), Bash(pip:*), Grep

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Initialize a chat model with explicit, version-safe defaults
  2. Iterate AIMessage.content as typed blocks, not strings
  3. Route across providers with a factory, not a conditional
  4. Count tokens correctly during streaming
  5. Pick the right with_structured_output method

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(python:*)
    • Bash(pip:*)
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • python.langchain.com
    • blog.langchain.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY
    • OPENAI_API_KEY
    • GOOGLE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Langchain Model Inference loads about 2.7k tokens when it runs, and up to ~7.4k if it reads all its reference files. Until then it costs about 121 tokens; SKILL.md has 784 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~121
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 784 words, ~2,683 tokens.

Download SKILL.mdSave it as .claude/skills/langchain-model-inference/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
langchain-model-inference
description
Invoke Claude, GPT-4o, and Gemini through LangChain 1.0 without tripping on the content-block, token-accounting, and structured-output quirks that silently break production code. Use when initializing chat models, routing across providers, iterating AIMessage content, or choosing a structured-output method. Trigger with "langchain model inference", "ChatAnthropic", "ChatOpenAI", "with_structured_output", "AIMessage content blocks", "langchain routing".
allowed-tools
Read, Write, Edit, Bash(python:*), Bash(pip:*), Grep
compatibility
Designed for Claude Code
version
2.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, langchain, langgraph, python, langchain-1.0, model-inference

LangChain Model Inference (Python)

Overview

AIMessage.content is a str on simple OpenAI calls and a list[dict] on Claude the instant any tool_use, thinking, or image block enters the response. Code that does message.content.lower() crashes with AttributeError: 'list' object has no attribute 'lower' — the #1 first-production-call LangChain 1.0 bug on Anthropic. And that is one of four separate "content shape" pitfalls in this skill:

  • P02 — AIMessage.content list-vs-string divergence
  • P03 — with_structured_output(method="function_calling") silently drops Optional[list[X]] fields on ~40% of real schemas
  • P05 — temperature=0 is not deterministic on Anthropic even though it is on OpenAI
  • P58 — Claude expects the system message at position 0; middleware that reorders messages makes it silently ignored

This skill walks through ChatAnthropic, ChatOpenAI, and ChatGoogleGenerativeAI initialization; model routing; token counting that is actually correct during streaming; content-block iteration; and a decision tree for with_structured_output methods that holds up on real schemas. Pin: langchain-core 1.0.x, langchain-anthropic 1.0.x, langchain-openai 1.0.x, langchain-google-genai 1.0.x. Pain-catalog anchors: P01, P02, P03, P04, P05, P53, P54, P58, P63, P64, P65.

Prerequisites

  • Python 3.10+
  • langchain-core >= 1.0, < 2.0
  • At least one provider package: pip install langchain-anthropic langchain-openai
  • Provider API key(s): ANTHROPIC_API_KEY, OPENAI_API_KEY, GOOGLE_API_KEY

Instructions

Step 1 — Initialize a chat model with explicit, version-safe defaults
python
from langchain_anthropic import ChatAnthropic
from langchain_openai import ChatOpenAI

claude = ChatAnthropic(
    model="claude-sonnet-4-6",
    temperature=0,
    max_tokens=4096,
    timeout=30,       # seconds. Default is None — hangs forever on provider stall.
    max_retries=2,    # Retries, not attempts. See P30 in pain catalog.
)

gpt4o = ChatOpenAI(
    model="gpt-4o",
    temperature=0,
    timeout=30,
    max_retries=2,
)

Explicit timeout and max_retries are not optional in production — the defaults are wrong for every workload we have measured. max_retries=6 (the default on ChatOpenAI) means a single logical call can bill as 7 requests on flaky networks.

Step 2 — Iterate AIMessage.content as typed blocks, not strings
python
from langchain_core.messages import AIMessage

def extract_text(msg: AIMessage) -> str:
    """Safe on both provider shapes. Works for streaming deltas too.

    Handles both dict blocks (provider-native) and typed block objects
    (LangChain 1.0 wrappers) — which Gemini, OpenAI tools, and future
    SDK versions may return.
    """
    if isinstance(msg.content, str):
        return msg.content
    parts = []
    for block in msg.content:
        # Block may be a dict (provider-native) or a typed object (1.0 wrapper)
        block_type = block.get("type") if isinstance(block, dict) else getattr(block, "type", None)
        if block_type == "text":
            parts.append(block["text"] if isinstance(block, dict) else block.text)
    return "".join(parts)

AIMessage.text() (1.0+) does this for you in most cases — prefer it. Roll your own only when you need to filter by block type (tool_use, image, thinking). See Content Blocks for the full block-type reference and streaming-delta shape.

Step 3 — Route across providers with a factory, not a conditional
python
from langchain_core.language_models import BaseChatModel

# Version-safe defaults applied to every model the factory builds.
# Callers can override via **kwargs.
_SAFE_DEFAULTS = {"timeout": 30, "max_retries": 2}

def chat_model(provider: str, **kwargs) -> BaseChatModel:
    defaults = {**_SAFE_DEFAULTS, **kwargs}  # caller's kwargs win
    if provider == "anthropic":
        return ChatAnthropic(model="claude-sonnet-4-6", **defaults)
    if provider == "openai":
        return ChatOpenAI(model="gpt-4o", **defaults)
    if provider == "gemini":
        from langchain_google_genai import ChatGoogleGenerativeAI
        return ChatGoogleGenerativeAI(model="gemini-2.5-pro", **defaults)
    raise ValueError(f"Unknown provider: {provider!r}")

A factory centralizes the version-safe defaults from Step 1 (timeout=30, max_retries=2) and the structured-output method pick from Step 5. Chains depend on the BaseChatModel protocol, not the concrete class. Callers override with chat_model("openai", timeout=60) when they need it.

Step 4 — Count tokens correctly during streaming

ChatAnthropic.stream() does not populate response_metadata["token_usage"] until the stream closes (P01). If your cost dashboard reads on_llm_end, it lags by the stream duration. Use astream_events(version="v2"):

python
async for event in claude.astream_events({"input": "..."}, version="v2"):
    if event["event"] == "on_chat_model_stream":
        chunk = event["data"]["chunk"]
        if hasattr(chunk, "usage_metadata") and chunk.usage_metadata:
            meter.record(chunk.usage_metadata["input_tokens"],
                         chunk.usage_metadata["output_tokens"])

See Token Accounting for per-provider differences (Anthropic reports input/output/cache separately; OpenAI aggregates; Gemini reports completion-only on stream start).

Step 5 — Pick the right with_structured_output method
ProviderModel classRecommended methodWhy
AnthropicClaude 3.5+, 4.xjson_schemaProvider-enforced, supports $ref and unions
OpenAIGPT-4o, GPT-4-turbojson_schemaStrict schema, additionalProperties: false enforced
OpenAIGPT-3.5, legacyfunction_callingPre-json_schema fallback
GeminiGemini 2.5 Pro/Flashjson_schemaNative structured output in 1.0+
AnyOlder or unknownjson_mode + Pydantic validate + retryJSON-parseable only, no schema enforcement (P54)
python
from pydantic import BaseModel, ConfigDict

class Plan(BaseModel):
    model_config = ConfigDict(extra="ignore")  # P53 — models add helpful extra fields
    steps: list[str]
    estimated_minutes: int

structured = claude.with_structured_output(Plan, method="json_schema")
plan = structured.invoke("Plan a 3-step deploy")

Avoid Optional[list[X]] fields — they silently return None on some providers (P03). See Structured Output Methods for a concrete comparison matrix and fallback pattern.

Show full SKILL.md (322 more words)Show less

Output

  • Chat models initialized with explicit timeouts (30s) and max_retries=2
  • Content-safe extractor that handles both str and list[dict] shapes
  • Factory-based routing with a single BaseChatModel return type
  • Streaming token counter that reports incrementally, not at stream end
  • with_structured_output chosen per provider capability, with Pydantic validation

Error Handling

ErrorCauseFix
AttributeError: 'list' object has no attribute 'lower'Treating Claude AIMessage.content as str (P02)Use msg.text() or the Step 2 extractor
ValidationError: extra fields not permittedPydantic v2 strict default; model added fields (P53)Set model_config = ConfigDict(extra="ignore")
ValidationError: Field required on Optional[list[X]]method="function_calling" drops ambiguous unions (P03)Switch to method="json_schema"
anthropic.BadRequestError: tool_choice requires toolsForcing tool without binding any (P63)Call .bind_tools([tool]) before .with_config(tool_choice=...)
google.api_core.exceptions.InvalidArgument: finish_reason=SAFETYGemini default safety thresholds (P65)Override safety_settings per model init or switch provider
Streaming response response_metadata["token_usage"] == {}Stream end not yet reached (P01)Use astream_events(version="v2")
ImportError: cannot import name 'ChatOpenAI' from 'langchain.chat_models'Legacy 0.2 import path (P38)from langchain_openai import ChatOpenAI

Examples

Routing: cheap draft, expensive final

A common pattern — draft with gpt-4o-mini, finalize with claude-sonnet-4-6. The factory in Step 3 makes this trivial; combined with with_structured_output the finalize step returns a typed object.

See Provider Quirks for the full draft-then-finalize example including the token budget calculation.

Extracting tool calls from a single-shot response

A classification task that should return one tool call with a typed argument. Use bind_tools([...], tool_choice={"type": "tool", "name": "Classify"}) for a single forced call — but never loop on a forced choice (P63).

See Structured Output Methods for the worked example and the decision tree for tool vs structured-output for extraction.

Multi-modal: screenshot plus prompt

Images are passed as content blocks, but the block shape differs between providers (P64). LangChain 1.0 abstracts this into a universal image content block.

See Content Blocks for the universal shape and per-provider adapter examples.

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/.curated/langchain-model-inference of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/content-blocks.md
  • references/one-pager.md
  • references/provider-quirks.md
  • references/structured-output-methods.md
  • references/token-accounting.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Langchain Model Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langchain Model Inference compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langchain Model Inference this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2.7kAutomated safety check: PassMIT
Tool Designagentailor/fullstack-langgraph-nextjs-agent132—~3.2kAutomated safety check: PassMIT
LangchainOrchestra-Research/AI-Research-SKILLs13k2 repos~3.2kAutomated safety check: PassMIT
Langchainlangchain-ai/docs426—~1.1kAutomated safety check: PassMIT
Sap Cloud SDK AIsecondsky/sap-skills462—~3.2kAutomated safety check: PassGPL-3.0
Scaffolding Openai Agentsaiskillstore/marketplace433—~3.3kAutomated safety check: PassNone

Similar skills

  • Tool Design

    agentailor/fullstack-langgraph-nextjs-agent

    Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).

    132 GitHub stars~3.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Langchain

    Orchestra-Research/AI-Research-SKILLs

    Framework for building LLM-powered applications with agents, chains, and RAG.

    13k GitHub starsUsed in 2 repos~3.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Langchain

    langchain-ai/docs

    Official

    Build agents with a prebuilt architecture and integrations for any model or tool.

    426 GitHub stars~1.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Sap Cloud SDK AI

    secondsky/sap-skills

    Integrates SAP Cloud SDK for AI into JavaScript/TypeScript and Java applications.

    462 GitHub stars~3.2k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Scaffolding Openai Agents

    aiskillstore/marketplace

    Builds AI agents using OpenAI Agents SDK with async/await patterns and multi-agent orchestration.

    433 GitHub stars~3.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Routerbase API Integration

    aiskillstore/marketplace

    Integrate applications with RouterBase, the OpenAI-compatible model gateway at https://routerbase.com/v1.

    433 GitHub stars~964 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Langchain Model Inference

What does Langchain Model Inference do?

Invoke Claude, GPT-4o, and Gemini through LangChain 1.0 without tripping on the content-block, token-accounting, and structured-output quirks that silently break production code. Langchain Model Inference is an agent skill from jeremylongshore/tons-of-skills-marketplace.0 without tripping on the content-block, token-accounting, and structured-output quirks that silently break production code.

When should I use Langchain Model Inference?

Langchain Model Inference fits situations like: initializing chat models; routing across providers; iterating AIMessage content; choosing a structured-output method.

How do I install Langchain Model Inference in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-model-inference -a claude-code`. Or copy the skill folder (skills/.curated/langchain-model-inference in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-model-inference in your project. Claude Code loads it when a task matches its description.

How do I install Langchain Model Inference in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-model-inference -a codex`. Or copy the skill folder (skills/.curated/langchain-model-inference in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-model-inference in your project. Codex loads it when a task matches its description.

Can I use Langchain Model Inference in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-model-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-model-inference, .gemini/skills/langchain-model-inference, .github/skills/langchain-model-inference and .opencode/skills/langchain-model-inference in your project.

What does Langchain Model Inference need to run?

Going by SKILL.md and its folder, Langchain Model Inference needs the command-line tools its instructions call (pip) and credentials named ANTHROPIC_API_KEY, OPENAI_API_KEY and GOOGLE_API_KEY. Our summary lists: Python 3; A credential in ANTHROPIC_API_KEY; A credential in OPENAI_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(python:*), Bash(pip:*), Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Langchain Model Inference access the network?

SKILL.md names 2 domains. As links in the text: python.langchain.com and blog.langchain.com. This is read from the text; nothing was executed.

Is Langchain Model Inference safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langchain Model Inference use?

Langchain Model Inference is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langchain Model Inference use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.7k tokens, read only when the agent opens those files.

What are the alternatives to Langchain Model Inference?

Skills that share tags, products or a category with Langchain Model Inference: Tool Design (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Langchain (Orchestra-Research/AI-Research-SKILLs, 13k stars), Langchain (langchain-ai/docs, 426 stars) and Sap Cloud SDK AI (secondsky/sap-skills, 462 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langchain Model Inference?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.