---
name: add-new-model
description: Add support for a newly-released language or image generation model in pydantic-ai (e.g. openai:gpt-5.6, anthropic:claude-sonnet-5, openai:gpt-image-2). Use when a provider ships a new model id and you need to wire literals, profile flags, adapters, and tests to recognize it. Handles SDK-lag, gateway list conventions, capability probing, and direct image-model geometry.
user-invocable: true
allowed-tools: Bash, Read, Edit, Write, Glob, Grep, WebFetch, WebSearch, AskUserQuestion
---

# Add New Model

Wire a newly-released provider model into pydantic-ai. Optimized for the common case (mirror an existing sibling); flags the cases where it's *not* a mirror and needs deeper work.

## Reference docs (read once before scoping)

- `agent_docs/pydantic-ai-slim.md` — the **Ownership** section, plus `pydantic_ai_slim/pydantic_ai/native_tools/AGENTS.md`, for the user-visible surface this model needs to land on.
- `pydantic_ai_slim/pydantic_ai/profiles/AGENTS.md`, `providers/AGENTS.md`, `models/AGENTS.md`, and `pydantic_ai_slim/pydantic_ai/AGENTS.md` (the capability-flag and `Provider.model_profile()` rules), plus the **Design Rules** section of `agent_docs/pydantic-ai-slim.md`. These tell you where capability facts belong (profile vs. provider vs. model class) when the new id has non-mirror behavior.

## Inputs

User invokes with `provider` + `model id` (e.g. `openai gpt-5.6`). If missing, ask via `AskUserQuestion`.

## Image generation models

Image-only models use a separate public surface from conversational models. If the model is consumed by `ImageGenerator`, update `KnownImageGenerationModelName` in `pydantic_ai_slim/pydantic_ai/images/__init__.py`, the relevant direct provider adapter, and its tests; do not also change conversational `KnownModelName`, profiles, gateway aliases, `ImageGenerationTool`, or `models/<provider>.py` unless that surface is explicitly supported and in scope. Add only the public model IDs the project intends to support, and do not infer or automatically add dated snapshots.

Keep common, provider-agnostic controls in `images/settings.py`, but import provider-specific setting types from the official SDK. Put model-specific size and aspect-ratio validation or mapping in the private `images/_<provider>_geometry.py` helper, and update the public support matrix in `docs/image-generation.md`. Verify geometry against official documentation; if the provider does not publish exact output shapes, probe every documented aspect-ratio and resolution combination for every supported model and record the evidence. Prefer deterministic table tests for the full matrix, adding one representative VCR cassette only when the new model or wire behavior needs integration coverage rather than recording every image combination.

## Step 1 — Verify the model exists at the provider

Never trust marketing names, news articles, or guesses. Hit the provider's model-listing endpoint:

| Provider  | Verification call |
|-----------|-------------------|
| OpenAI    | `curl -s https://api.openai.com/v1/models -H "Authorization: Bearer $OPENAI_API_KEY"` |
| Anthropic | `curl -s https://api.anthropic.com/v1/models -H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01"` |
| xAI       | `curl -s https://api.x.ai/v1/models -H "Authorization: Bearer $XAI_API_KEY"` |
| Google    | `curl -s https://generativelanguage.googleapis.com/v1beta/models -H "x-goog-api-key: $GOOGLE_API_KEY"` |
| Groq      | `curl -s https://api.groq.com/openai/v1/models -H "Authorization: Bearer $GROQ_API_KEY"` |
| Bedrock   | `aws bedrock list-foundation-models --region "$AWS_REGION"` |

Load credentials from the repo-root `.env` with `source .env && <cmd>`. `list-foundation-models` is region-scoped, so query the region your models are actually deployed in (not a hard-coded default). List **every** id the provider exposes for this release — base, dated snapshot, `-pro`, `-mini`, `-nano`, `-codex`, `-chat-latest`. Add only what actually exists; do not extrapolate sibling variants.

If the user-given id is **not** in the listing, stop and confirm with the user before proceeding.

## Step 2 — Mirror the most recent add-model PR for this provider

```bash
git log --all --oneline --grep="<previous-version-pattern>" -20
# e.g. for openai: --grep="gpt-5\.4\|gpt-5\.3"
# e.g. for anthropic: --grep="claude-opus-4\|claude-sonnet-4"
```

Pick the smallest, most recent "add model X" PR for the same provider. Pull its file list with `gh pr view <num> --json files --jq '.files[].path'`. **That file list is the floor of what you'll touch.** It is rarely the ceiling.

## Step 3 — Enumerate (load-bearing step)

For every variable, tuple, and literal you're about to touch, grep its readers across the repo. **This step is what catches the snapshot/enumeration tests that ratchet on every model add.** Skipping it pushes work onto CI and produces broken PRs.

Specifically, for a typical model add, grep for:

- The previous model id literal you're mirroring (e.g. `gpt-5.4`, `claude-opus-4-5`) — `rg '<prev-id>' --glob '!**/*.yaml' --glob '!**/cassettes/**'`
- Search for the previous model's display name in docs and public docstrings; inspect capability rosters for the same features.
- Every prefix/membership key in the profile module you're editing (e.g. OpenAI's `_REASONING_SUPPORT_BY_PREFIX` keys, Anthropic's inline `model_name.startswith((...))` tuples, xAI's `_GROK_43_REASONING_MODELS`)
- `KnownModelName` and its provider-block neighbours
- Snapshot test files: `tests/models/test_model_names.py`, `tests/test_capability_spec.py`

Classify each hit:
- **must update** — model-name lists, dispatch tuples
- **snapshot to refresh** — `inline_snapshot` blocks needing `pytest --inline-snapshot=fix`
- **skip** — VCR cassettes, docs about an unrelated model

If `rg` output looks mangled (unicode/regex artifacts), drop to `grep -n` — don't push past garbled output.

## Step 3b — Pair the genai-prices entry

Cost and `context_window` data come from `pydantic/genai-prices` through `_genai_prices.py`.
Bundled data requires a `genai-prices` release.
Update the Pydantic AI lock to verify the bundled entry locally.
The live merged feed can supply entries through `pydantic_ai.prices.update_in_background()`.
Check the updater for an entry before claiming a package release blocks support.
`Model.profile` only consults genai-prices when nothing else set `context_window`.
Without a price entry in either source, `ModelResponse.cost()` raises `LookupError` and a
`cost_limit` cannot be enforced — the run warns `CostNotFoundWarning` at the end instead.
Without a `context_window` value from any source, `RunContext.context_window_used` is `None`.
If no genai-prices entry exists, open the genai-prices PR alongside the model add and link the two.

Before you write the entry, check that no one has added it already. Someone else may have added
it on release day. Grep genai-prices `main` for each provider file you plan to edit, then the
provider-file changes in its open PRs:

```bash
for f in <provider files, e.g. openai openrouter github_copilot>; do gh api "repos/pydantic/genai-prices/contents/prices/providers/$f.yml" -H 'Accept: application/vnd.github.raw' | grep -n '<id>' | sed "s/^/$f.yml:/"; done
for n in $(gh pr list --repo pydantic/genai-prices --state open --limit 200 --json number --jq '.[].number'); do gh api "repos/pydantic/genai-prices/pulls/$n/files" --paginate --jq '.[] | select(.filename | startswith("prices/providers/")) | .patch' | grep -q '<id>' && echo "#$n"; done
```

A hit means you link that entry or PR instead of opening your own. Do not rely on `gh search prs`:
its index lags newly opened PRs, and on release day it did not yet return genai-prices #732.

Check the current catalogs of other providers that host the new model before scoping that PR.
For example, OpenRouter may publish `openai/<id>` and a `YYYYMMDD` canonical slug on release day
even when OpenAI's own model list exposes only the base id. Add a separate genai-prices entry
for each confirmed host with its own published rates; do not infer Bedrock or Azure availability
from an older sibling model.

Compare the new price entry's `match` with the adjacent model families in the same provider file,
and accept the same id forms they do, even before the provider publishes a snapshot. That way a
later snapshot gets the same price and context data. The forms differ per provider file:

| genai-prices file | Forms the siblings accept |
|---|---|
| `openai.yml` | base id + `-YYYY-MM-DD` |
| `anthropic.yml` | base id + `-YYYYMMDD` (Claude Opus entries also take `claude-opus-X.Y` / `claude-X-Y-opus` aliases) |
| `google.yml` (Vertex Claude) | base id + `@YYYYMMDD` |
| `aws.yml` (Bedrock) | `global.` and `us.`/`eu.`/`jp.`/`au.` profiles, bare `anthropic.` id, and `-v1` / `-v1:0` suffixes |
| `openrouter.yml` | slug + `:beta` |

Dump the siblings' clauses before writing yours rather than copying one PR's shape (see #8635 and
genai-prices #709). Test the base and suffixed forms, including the canonical model id. This match
rule does not add a speculative id to Pydantic AI's model-name literals. Its profile prefix should
resolve the same capabilities for either form.

## Step 4 — SDK pin check

Snapshot/enumeration tests in this repo often tie `KnownModelName` to a literal set defined in the provider SDK. **The provider SDK frequently lags the model release by days.**

For OpenAI, check the broad union the repo actually consumes (`OpenAIModelName = str | AllModels`), **not** the chat-only `ChatModel` Literal — `AllModels` also carries Responses-API-only and embeddings ids that the enumeration test walks:

```bash
uv run python -c "from openai.types import AllModels; from typing import get_args; print([m for m in get_args(AllModels) if '<new-version>' in m])"
```

If no SDK release lists the new id, bridge it with a local `Literal` on the model-name alias. The PR then lands green without waiting for the SDK.

If a release lists the id but sits inside the 7-day `exclude-newer` window, the providers differ:

- OpenAI: bridge anyway with `OpenAIModelName = str | AllModels | Literal['<id>']`. The docstring names the `openai` release that adds the id. A later PR bumps the floor and drops the bridge (#8635 bridged, #8655 dropped).
- Anthropic: bump the SDK through the quarantine exemption its landmine section describes. Anthropic checks `ModelParam`, not `Model`.
- xAI: see the SDK-lag bridge notes in its landmine section.

## Step 5 — Probe capabilities (only if not a pure mirror)

If the new model is just another sibling in an existing family (e.g. `gpt-5.5` after `gpt-5.4`), skip to Step 6 — the existing profile branch covers it once you add the prefix to the dispatch tuple.

If the model is a new family or has unclear capabilities, write a small comparison script (`local-notes/probe_<model>.py`) that hits the new model AND its closest neighbour with:
- `temperature` / `top_p` (does the API reject sampling params?)
- `reasoning.effort` values (`none`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`) — note which the API accepts. OpenAI's 400 lists the accepted values
- New parameters mentioned in the release notes
- Streaming / tool calls if the family is new

Diff the responses. Anything that diverges from the neighbour belongs in the profile.

### Gateway parity

Where the gateway serves a model, it must behave the same as the provider's canonical API. Step 3 only
gets the id recognized; this is about behavior, and nothing enumerates it for you.

The gateway reaches the canonical API through an ordinary SDK client carrying a proxy base URL. So:

- **Narrow a capability by client class, never by base URL.** Bedrock, Vertex and Foundry are separate
  transports and earn their own gates. A proxied client is the canonical API, and must keep every
  capability the unproxied one has.
- **A `base_url` test inside a capability decision is the defect, not the fix.** It splits the gateway
  off from the transport it actually reaches. No capability in `models/` or `profiles/` is decided that
  way — if you are about to be the first, you are answering the wrong question.
- **Probe the gateway leg rather than reasoning about it.** `Model('<id>', provider='gateway')`, then
  exercise whatever capability you gated. If `PYDANTIC_AI_GATEWAY_BASE_URL` is set in the environment,
  check it points at the gateway root: a provider-specific proxy path 404s every other provider.
- **Probe Bedrock Gateway profile regions individually.** A successful `gateway/bedrock:` inference-profile ID
  does not establish support for another region prefix or the bare ID. Probe each candidate; exclude only IDs
  Gateway rejects.

A model the gateway genuinely does not serve is the other case entirely: it belongs in
`UNSUPPORTED_GATEWAY_MODEL_NAMES`, on evidence that the gateway rejects the id. Never leave the id
advertised and quietly degraded by a capability carve-out instead.

## Step 6 — Edit (minimal diff matching the mirrored PR)

Make only the changes the enumeration step surfaced. Resist scope creep. If you suspect a pre-existing bug in a sibling model's profile, reproduce it live first. Fold a reproduced bug on the same profile gate into this PR, with a test and one PR-body line. Flag an unreproduced or larger one in the PR description instead.

After edits:

```bash
make format && make lint
PYRIGHT_PYTHON_IGNORE_WARNINGS=1 uv run pyright <changed-python-files>
```

Run the tests directly touching the changed surface — the profile test plus any enumeration tests you updated. CI is the safety net for the long tail; locally you only need to verify the surface area of your change.

If snapshot tests changed: `uv run pytest <file> --inline-snapshot=fix` then verify the diff is the expected literal addition only.

## Step 7 — VCR / integration tests

**Default for mirror-only adds:** skip recording a new VCR. Repo convention uses one representative model per family for VCR (e.g. `gpt-5.2` covers the gpt-5.x reasoning family). The profile unit test added in Step 6 is sufficient.

**When the new model introduces meaningful changes** to `pydantic_ai_slim/pydantic_ai/models/<provider>.py` (new request shape, new response field, new handler branch):

1. Look for an existing parametrized VCR test that covers the changed feature. `rg -l '<feature-name>' tests/models/`. If one exists and it parametrizes over model ids, **tag the new id onto the parametrize list** rather than writing a new test.
2. If no parametrized coverage exists and you need a new VCR test, place it:
   - **Prefer** `tests/models/<provider>/test_<feature>.py` **only if the file already exists** (e.g. `tests/models/anthropic/test_output.py`).
   - Otherwise add it to `tests/models/test_<provider>.py`. **Do not create a new `tests/models/<provider>/` subdirectory** if one doesn't already exist for this provider.
3. Record using the `testing-skill` skill workflow.

## Step 8 — PR

Follow the `pushing-commits-to-the-repo` skill for the title, body, template, and final metadata
check. Keep the model-specific evidence concise:

- One sentence: what model(s) were added.
- "Verified via probe / mirror of #NNNN" — explicit about which changes were API-verified vs assumed-by-mirror.
- Flag pre-existing latent bugs found but deliberately not fixed.
- Link the prior add-model PR for context.

## Provider-specific landmines

### OpenAI

- **`_REASONING_SUPPORT_BY_PREFIX`** in `pydantic_ai_slim/pydantic_ai/profiles/openai.py` — a dict keyed by model-name prefix (`'gpt-5.6'`, `'gpt-5.3-chat'`, `'gpt-5'`, `'o'`, …) → `_ReasoningSupport(enabled_by_default, can_be_disabled, supports_mode, supports_context)`, resolved **first-match-wins** by `_reasoning_support()`. A new `gpt-5.N` family MUST be added here, and **ordering matters**: a more specific prefix (`'gpt-5.3-chat'`) must precede the broader one it would otherwise shadow (`'gpt-5.3'`), and every newer `gpt-5.x` family must precede the plain `'gpt-5'` catch-all. Miss it and the model falls through to the `_NO_REASONING` default (`thinking_always_enabled=False`, `openai_supports_reasoning_effort_none=False`) — wrong defaults, no error. The resolved matrix is pinned in `tests/profiles/test_openai.py`.
- **`KnownModelName` lives in `pydantic_ai_slim/pydantic_ai/models/_known_model_names.py`** (a `TypeAliasType`), **not** `models/__init__.py`. It has split `openai:` and `gateway/openai:` blocks. Don't assume the gateway block omits `-pro`/`-chat-latest` — for the `gpt-5.x` series it enumerates them (`gateway/openai:gpt-5.2-pro`, `gateway/openai:gpt-5.3-chat-latest`, …). Mirror the exact enumeration of the most recent series across both blocks rather than guessing a convention.
- **Most `gpt-5.x-chat` variants DO reason** (`_ALWAYS_ON_REASONING`: reason at a fixed effort, reject `reasoning_effort='none'` and sampling parameters). The non-reasoning exception is the original `gpt-5-chat`/`gpt-5-chat-latest` (`_NO_REASONING`). Verify each `-chat`/`-chat-latest` variant against the live Responses API; don't copy a sibling's reasoning class blindly.
- **`-pro` variants map to `_ALWAYS_ON_REASONING`** (`gpt-5.2-pro`, `gpt-5.4-pro`, `gpt-5.5-pro`) — they reason and reject `effort='none'`. `_ReasoningSupport` doesn't encode per-effort-*value* rejection, so if a new `-pro` rejects a specific value (e.g. `'low'`), flag it rather than assuming the enum covers it.
- **`tests/models/test_model_names.py::test_known_model_names`** asserts `known_model_names()` equals the set generated from `_PROVIDER_TO_MODEL_NAMES['openai']`, i.e. `OpenAIModelName = str | AllModels` (the broad union, not the chat-only `ChatModel`). A literal missing from `AllModels` fails this test — Step 4's SDK check is mandatory and must query `AllModels`.
- **`tests/test_capability_spec.py::test_model_json_schema_with_capabilities`** is a snapshot test enumerating every `KnownModelName`. Refresh with `--inline-snapshot=fix`.
- **A dotted point release matches none of its family's prefixes.** `gpt-6.1-sol` does not start with `'gpt-6-sol'`. Before you add it, it resolves to `_NO_REASONING` and misses every `startswith` gate in `openai_model_profile()`. Add it to `_REASONING_SUPPORT_BY_PREFIX` and to the generation's gate tuple (`_GPT_6_MODEL_PREFIXES`).
- **Probe a point release against every sibling, not only its namesake.** GPT-6.1 Sol rejects `effort='none'` like GPT-6 Astra; GPT-6 Sol accepts it. The accepted-effort list decides `can_be_disabled`. Decide `supports_image_output` by forcing the tool (`tool_choice={'type': 'image_generation'}`), not from the model page's tool list.
- **Chat Completions rejects function tools while reasoning is on** for the GPT-6 family: `Function tools with reasoning_effort are not supported`. A model that rejects `effort='none'` therefore has no Chat Completions tool calling. Document the limit in `docs/models/openai.md`.
- **A gateway 404 `No cost data available for model` means genai-prices has no entry yet.** It is not a gateway rejection. Keep the `gateway/openai:` literal and leave `UNSUPPORTED_GATEWAY_MODEL_NAMES` alone. The id waits on the Step 3b genai-prices entry.
- **clai2 keeps a curated Codex menu**: `CODEX_MODELS` in `src/pydantic_clai2/pydantic_clai2/model_catalog.py`. Add the id when Codex offers it. Check `openai/codex`'s `codex_tui__chatwidget__tests__model_selection_popup.snap`, then run `Agent('openai-codex:<id>')` with a local Codex login. Update the Codex model lists in `src/pydantic_clai2/README.md` (two of them), `src/pydantic_clai2/PLUGINS.md` and `docs/harness/clai2.md`: `rg 'openai-codex:gpt-|gpt-5.6-luna' src/pydantic_clai2 docs/harness`.

### Anthropic

- **TWO literal lists, not one.** Add the id to BOTH:
  1. `pydantic_ai_slim/pydantic_ai/models/_known_model_names.py` — the `anthropic:` AND `gateway/anthropic:` blocks (the `KnownModelName` alias moved here from `models/__init__.py` in #5803; older add-model PR diffs that edit `__init__.py` are stale on this point).
  2. `AnthropicModelName` in `models/anthropic.py` — see the SDK-lag bridge below.
- **Anthropic names ARE enumeration-tested**, unlike what you might assume from the hand-maintained look of the list. `tests/models/test_model_names.py::test_known_model_names` asserts `known_model_names()` (i.e. `KnownModelName`) equals the set generated from `_PROVIDER_TO_MODEL_NAMES['anthropic']`, which is `AnthropicModelName` = `ModelParam` (from the installed `anthropic` SDK) `| Literal[...bridge...]`. A new id missing from BOTH the SDK's `ModelParam` and the local bridge fails this test with `Extra names: {...}`.
- **SDK-lag bridge (the Step 4 mechanism for Anthropic).** When the installed SDK's `anthropic.types.model_param.ModelParam` doesn't yet list the new id (check: `get_args` it and grep), bridge it with a local `Literal`:
  ```python
  AnthropicModelName = LatestAnthropicModelNames | Literal['claude-fable-5']
  ```
  plus a docstring note to drop the literal once the `anthropic` pin is bumped past the release that adds it. This is the in-repo pattern (commit `87e7ccf39`, PR #5849, added the `claude-fable-5` bridge; `526b065e2` later dropped it and bumped the floor to `anthropic>=0.108.0`). The bridge lands green immediately — no need to split the PR for Anthropic. NOTE: `ModelParam` ≠ `anthropic.types.model.Model`; check `ModelParam` (it's the superset the repo actually consumes, and may carry ids `Model` doesn't).
- **Capability flags live as `startswith` prefix tuples in `profiles/anthropic.py`** inside `anthropic_model_profile()` (+ the module-level `_ANTHROPIC_CODE_EXECUTION_20260120_MODEL_PREFIXES`). A new family is NOT a literal-only add — it almost always needs at least one profile override (a literal-only add is only right when the family truly inherits every default branch, which is rare). Probe and set each independently: `models_that_support_json_schema_output`, `supports_adaptive`, `supports_effort`, `supports_xhigh_effort`, `disallows_budget_thinking`, `disallows_sampling_settings`, `supports_task_budgets`, `supports_tool_search`, code-exec version, `anthropic_supports_fast_speed`. Default-`False` flags (e.g. fast speed) are subtractive — just omit the id from that tuple.
- **A point release inherits every flag of its base id silently.** The tuples are `startswith` prefixes, so `'claude-opus-5'` already matches `claude-opus-5-5` (as `'claude-fable-5'` matches `claude-fable-5-1`): before you touch anything, the new id resolves to the base model's profile. Tests stay green and nothing warns, so the only way to find a divergence is to read the model's migration guide and probe side by side with the base id. Opus 5.5 looked like an Opus 5 mirror and broke default `output_type` runs with a 400 until it opted out of forcing. Where a flag must *not* carry over, carve the id out explicitly (`startswith('claude-opus-5') and not startswith('claude-opus-5-5')`).
- **Read the migration guide's "breaking changes" before probing.** Anthropic's `platform.claude.com/docs/en/models/<id>/migration-guide` and `whats-new-<id>` pages list every divergence from the previous model and name which other models share it (e.g. "the first three also apply on Claude Fable 5.1"). Those map straight onto profile flags, and they tell you what to probe.
- **Bump the SDK through the 7-day quarantine rather than bridging, once the SDK lists the id.** `exclude-newer = "7 days"` in the root `pyproject.toml` keeps a same-day `anthropic` release out of the lock; admit exactly that release with a timestamp cutoff under `[tool.uv.exclude-newer-package]` (one second past its last artifact's PyPI `upload_time`, plus a `TODO` to remove it once the global window covers it — past that date it turns into a ceiling), raise the floor under `tool.hatch.metadata.hooks.uv-dynamic-versioning.optional-dependencies` in `pydantic_ai_slim/pyproject.toml` (move the exact `uv add --optional`-generated requirement there and remove its conflicting `[project.optional-dependencies]` table), run `uv lock --upgrade-package anthropic`, and refresh the gh-aw runner's own lock with `uv lock --script .github/scripts/pydantic-ai-runner` (no CI step checks it, so a stale one stays green while silently dropping its pinned hashes). File a tracking issue for removing the cutoff and link it from the `TODO`. Precedents: #7989 (1.3.0), #8637 (1.8.0). The local-`Literal` bridge below is the fallback for when no SDK release lists the id yet.
- **An SDK bump is a real change: pyright `models/anthropic.py` and the Anthropic tests against it.** 1.8.0 renamed the citations request TypedDict to `BetaCitationsConfigParamParam` (`BetaCitationsConfigParam` became a response model, and passing it into a request param broke a dict assertion) and widened `BetaInputTransformation` to a union with `thinking_mismatch_allowed`.
- **Opus 5.5 skips thinking on trivial prompts at its default `medium` effort.** A cassette test that needs a thinking block (e.g. `stale_thinking_block_history`) has to raise `anthropic_effort` for it.
- **A point release falls into its base model's price entry.** The base entries' prefix and `contains` clauses (`starts_with: claude-opus-5`) also capture `claude-opus-5-5`. So until the new entry exists, `calc_price` returns the *old* model's price with no error: genai-prices 0.1.7 priced Opus 5.5 at Opus 5's $5/$25 instead of $4/$20. Check `calc_price(..., model_ref='<new-id>')` against the published price. The genai-prices PR has to narrow the base entry's clauses so they stop at the base model, keeping every form they matched before and pinning those forms with a positive test (genai-prices #671, #709), as well as add the new entry per Step 3b.
- **Forced `tool_choice` is a real per-model divergence worth probing.** Most Anthropic models accept `tool_choice` `{'type':'any'}`/`{'type':'tool'}` and only reject forcing alongside *thinking*; Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5.5 reject it **unconditionally** (400 `tool_choice forces tool use is not compatible with this model`). That's modeled by `AnthropicModelProfile.anthropic_supports_forced_tool_choice` (default `True`) threaded into `_support_tool_forcing` in `models/anthropic.py`. Probe `tool_choice={'type':'any'}` against the new id AND its neighbour to tell a genuine divergence from a thinking-only constraint.
- **Probe the stale-thinking-block retry shape on every new binding model.** When the request set no thinking, the retry sends a `thinking` object through `extra_body`. Claude Fable 5.1 and Claude Opus 5.5 accept `{'block_binding': ...}` alone; Claude Sonnet 5.5 rejects it with `thinking.type: Field required`, so the retry fills in `'type': 'adaptive'`. Probe `extra_body={'thinking': {'block_binding': {'prefix_mismatch_behavior': 'drop_block'}}}` against the new id. Also probe any new `thinking.type` the release adds with `block_binding`: Sonnet 5.5's `between_tools` rejects it (`Extra inputs are not permitted`), so the retry skips that type.
- **Probe a mid-conversation `system` entry against the `<system>`-tagged fallback before touching `_INLINE_SYSTEM_PROMPT_MODEL_PREFIXES`.** Ask each rendering to lift a restriction the top-level prompt set; a plain non-conflicting instruction lands on every model and discriminates nothing. Claude Opus 5 obeys the entry 6/6 and the tagged text 0/6. Claude Sonnet 5.5 refuses both 6/6, so Anthropic's published support decides and it stays in: the entry keeps operator authority at no measured cost. Leave an id out only when the tagged fallback does measurably better, as it did on Claude Sonnet 5. `_TOOL_AVAILABILITY_DELTA_MODEL_PREFIXES` is a separate gate: probe a `tool_addition` by reference and check that the model calls the added tool.
- **List Bedrock inference profiles, not only foundation models.** Claude Sonnet 5.5 launched with `global.anthropic.claude-sonnet-5-5` and no `us.` profile. Run `aws bedrock list-inference-profiles` and add only the profiles it returns.
- **Tests:** profile-flag unit tests go in `tests/profiles/test_anthropic.py` (NOT `tests/models/test_anthropic.py`). Forced-tool-choice / `_prepare_tools_and_tool_choice` fallback tests go in `tests/models/test_tool_choice_unit.py`. The capability behaviors keyed on shared flags (sampling drop, budget-thinking reject, xhigh) are already covered by the opus-4-7/4-8 parametrized tests — adding the new id to those lists is redundant once a dedicated profile test asserts the flags.
- **`tests/test_capability_spec.py::test_model_json_schema_with_capabilities`** snapshots the whole `KnownModelName` enum. Refresh it by running THAT TEST ALONE with `--inline-snapshot=fix` — running the whole file can pull in unrelated `snapshot()` blocks and abort the fix.
- **`providers/bedrock.py` `bedrock_structured_output_unsupported`**: only relevant if the new id is actually served on Bedrock. A direct-API-only model (not in Bedrock's foundation-model list) doesn't belong there; don't add it speculatively just because the mirrored PR did.

### xAI (Grok)

- **Strict enumeration despite `XaiModelName = str | ChatModel`.** The `str` arm looks permissive but the enumeration test's `get_model_names` recurses into the union and yields nothing for a bare `str` type — so `KnownModelName`'s `xai:` block is strictly enforced against the SDK's `ChatModel` Literal, exactly like OpenAI. `tests/models/test_model_names.py::test_known_model_names` fails with "Extra/Missing names" on any mismatch. Confirm parity: `xai:` + `get_args(ChatModel)` must equal the `xai:` entries in `models/_known_model_names.py`.
- **SDK-lag bridge (Anthropic-style, and it's needed for xAI too).** `xai_sdk`'s `ChatModel` frequently lags a release — as of 1.17.0 it still lacked `grok-4.5`, so **bumping the floor won't help** (check newer wheels first: download from PyPI and grep `xai_sdk/types/model.py` for `ChatModel: TypeAlias = Literal[`). Bridge with a local Literal: `XaiModelName = str | ChatModel | Literal['grok-4.5', 'grok-4.5-latest']`, docstring-note to drop it when the floor is bumped past the release that adds the id. This makes the enumeration test's generated side include the new id, matching the hand-added `_known_model_names.py` literal — lands green immediately. (Historically xAI *bumped the SDK floor* — commits `e3f6e3c54`/`58f394aea` — but that only works when the SDK already ships the id.)
- **A new `grok-4.x` is NOT a pure mirror.** Reasoning-effort support lives in `profiles/grok.py` as membership sets (`_GROK_43_REASONING_MODELS` + a per-family effort frozenset), not startswith prefixes. The `grok-4` prefix auto-grants `grok_supports_builtin_tools=True` but leaves `grok_reasoning_efforts` **empty** (→ `supports_thinking=False`) unless you add the id to a reasoning-models set. Forgetting this silently ships a reasoning model with thinking off. Add a `_GROK_<ver>_REASONING_MODELS` set + effort frozenset and an `elif` branch in `grok_model_profile`.
- **Probe reasoning efforts via the OpenAI-compatible REST endpoint**, comparing against the closest neighbour: `POST https://api.x.ai/v1/chat/completions` with `{"model":..., "reasoning_effort": <val>, "max_tokens":1}`. A rejected value returns 400 `This model does not support 'reasoning_effort' value '<val>'`. **Whether `none` is accepted decides `thinking_always_enabled`** (rejected → always-on). CAVEAT: REST silently accepts `xhigh`/`minimal` even though the gRPC `ReasoningEffort` (in `xai_sdk/types/chat.py`) is `Literal['none','low','medium','high']` — don't over-read REST acceptance; `GrokReasoningEffort` is those four and `_map_reasoning_effort` collapses `xhigh`→`high`, `minimal`→`low`. Grok 4.5 example: accepts `low/medium/high`, rejects `none` → always-on; Grok 4.3 accepts `none` too.
- **Floating aliases** (`grok-latest`, `grok-build-latest`) go in the profile reasoning-models set (so passing them resolves the right behavior) but are **NOT** added as `KnownModelName` literals — mirror the SDK, which lists only stable ids like `grok-4.3`/`grok-4.3-latest`.
- **xai is NOT a gateway provider** (`'xai'` absent from `providers/gateway.py`'s `ModelProvider`) — no `gateway/xai:` entries in `_known_model_names.py`.
- **Snapshot that ratchets:** `tests/test_capability_spec.py::test_model_json_schema_with_capabilities` embeds the full `KnownModelName` enum. It's a plain sorted string list — hand-add the new ids in sorted position (deterministic, no need for `--inline-snapshot=fix`). Profile-flag tests go in `tests/providers/test_xai.py` (see `test_xai_model_profile`); the parametrized `tests/test_thinking.py::test_grok_43_profile_thinking_support` asserts the *4.3* effort set specifically — don't add a different-effort model to it.
- **env / probing:** `XAI_API_KEY` lives in the repo-root `.env` (not in every worktree). Run probes with `source .env && <script>` so `$XAI_API_KEY` is exported; put any `curl` referencing it in a script file rather than passing the key inline. Verify enumeration/profile logic with a plain `uv run python` snippet (recurse `get_args(XaiModelName)`, compare to `known_model_names()`; call `grok_model_profile(...)` directly) rather than a full `uv run pytest tests/` run.

### Bedrock

- **Bedrock Mantle is a separate provider from Bedrock Runtime.** `bedrock:` (the `BedrockProvider`, boto3-only) talks to the Converse API; `bedrock-mantle:` (the `BedrockMantleProvider`, an `openai`-backed `Provider[AsyncOpenAI]` built on `AsyncBedrockOpenAI`) talks to Mantle's OpenAI-compatible API. They have separate model catalogs and separate optional extras (`bedrock` vs `bedrock-mantle`); don't fold Mantle deps into the `bedrock` group.
- **Mantle model families use different endpoints, keyed off the profile.** `BedrockMantleProvider.model_profile` stamps `bedrock_mantle_interface: Literal['chat','responses','openai-responses']` on the profile (GPT-5.4+ → `openai-responses` at `/openai/v1`; GPT-OSS → `responses` at `/v1`; GPT-OSS Safeguard → `chat` at `/v1`). `infer_model` reads that (via the profile, not a separate interface method) to pick `BedrockMantleResponsesModel` vs `BedrockMantleChatModel`, and the Responses model overrides `client` to pick the base URL. Add a family only after verifying its endpoint against the AWS model card + a live request.
- **`bedrock:` stays on Converse; it does NOT auto-route to Mantle.** A GPT-5.4+ model on `bedrock:` raises from `BedrockProvider.model_profile` pointing users to `bedrock-mantle:` (there's a `TODO(v3)` to flip the default with a deprecation later). Only add `bedrock-mantle:` names to `KnownModelName` — no `bedrock:openai.gpt-5.*` names, and hence no `UNSUPPORTED_GATEWAY_MODEL_NAMES` entries for them.
- **Response-scoped tool-call IDs are a profile flag, not a Mantle-wide behavior.** `openai_responses_tool_call_ids_are_response_scoped` (on `OpenAIModelProfile`) is enabled only for Mantle GPT-5.6 Responses; `OpenAIResponsesModel` qualifies call IDs with the response ID in both request and streaming ingestion so history stays uniquely keyed (#6536).

### Google (Gemini)

- **TWO places for the id, FOUR `KnownModelName` blocks.** Add to:
  1. `LatestGoogleModelNames` in `models/google.py` (`GoogleModelName = str | LatestGoogleModelNames` — the `str` arm is permissive at typecheck time, but the enumeration test only walks the `Literal` arm).
  2. `models/_known_model_names.py` — **four** blocks: `gateway/google-cloud:`, `gateway/google:`, `google-cloud:`, `google:` (older add-model PRs that only edit three blocks or `models/__init__.py` are stale; KnownModelName moved in #5803).
- **No SDK-lag bridge needed.** `google-genai` does not ship a model-id Literal the enumeration test consumes — the local `LatestGoogleModelNames` Literal *is* the source of truth. Adding the id lands green immediately.
- **Profile is substring-gated, with a per-model level table.** `profiles/google.py` keys off `'gemini-3' in model_name` (thinking level, tool combination, server-side tool invocations, MIME types in tool returns) and `'pro' in model_name and 'flash' not in model_name` (always-on thinking). The exception is `_MODEL_THINKING_LEVELS`, a `startswith` table mapping id prefixes to their documented level sets that already holds both pro previews, the 3.7 and 3.8 flash ids, and `gemini-3.1-flash-lite-image` — so probe every new id rather than assuming the Gemini-3 branch covers it. Probe all four levels with `generateContent` and `thinkingConfig.thinkingLevel` (`MINIMAL`, `LOW`, `MEDIUM`, `HIGH`) on the Gemini API and on Vertex separately. A 400 on both means the id needs an entry in the table carrying exactly the levels it accepts (non-contiguous sets like `minimal, high` are fine — unsupported unified efforts snap to the nearest documented level). A 400 on only one of them means a level set in `GoogleModel.profile` gated on the client's transport (`_is_google_cloud`), as `gemini-3.1-flash-image` has for the Gemini API, so the other API keeps the levels it accepts. Reach Vertex with application-default credentials, `GOOGLE_PROJECT`, and `location='global'`, as `tests/conftest.py::vertex_provider` does. If you can probe only the Gemini API and it 400s, use the transport-gated branch too, so Vertex keeps its current behavior. Cite Google's documented level sets in the code comment: the Gemini API [thinking table](https://ai.google.dev/gemini-api/docs/thinking) (for image models, the [image-generation page](https://ai.google.dev/gemini-api/docs/image-generation) instead) and the Vertex [thinking table](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/thinking). The probe results decide the entry, even where they diverge from those tables. Probe too when the release notes claim any other capability divergence (no thinking, image-only, Pro always-on).
- **API verification:** `curl -s "https://generativelanguage.googleapis.com/v1beta/models?pageSize=200&key=$GOOGLE_API_KEY"` (key is often in the main worktree `.env`, not every linked worktree). Confirm exact ids; do **not** invent dated snapshots or `-preview` suffixes. Specialized / limited-access models (e.g. Flash Cyber via CodeMender) are out of scope unless they appear in that public listing.
- **Gateway support is opt-out, not opt-in.** The enumeration test generates `gateway/{google,google-cloud}:*` for every `LatestGoogleModelNames` entry **except** those listed in `UNSUPPORTED_GATEWAY_MODEL_NAMES` in `tests/models/test_model_names.py`. Mirror the most recent sibling series: if `gemini-3.5-flash` is in the gateway KnownModelName blocks (not in the unsupported set), new flash siblings go there too. Only add to `UNSUPPORTED_GATEWAY_MODEL_NAMES` when the gateway actually rejects the id.
- **Snapshots / tests:** hand-add the new ids in sorted position in `tests/test_capability_spec.py::test_model_json_schema_with_capabilities` (plain sorted string list). Mirror-only adds skip new VCR by default; #5527 recorded one for `gemini-3.5-flash` but that is not required for a pure name add.
- **Docs:** example snippets often hard-code a recent flash id (`docs/models/google.md`, `docs/capabilities/thinking.md`) — leave them alone unless the docs maintain a model registry table (they currently do not).

Google image-model landmines:

- Direct image generation has a separate public literal, `KnownImageGenerationModelName` in `pydantic_ai_slim/pydantic_ai/images/__init__.py`. When the task is scoped to `ImageGenerator`, update and test this literal independently; do not automatically widen the change to conversational `KnownModelName`, gateway aliases, profiles, and capability snapshots unless those surfaces are explicitly in scope.
- `Client().models.list()` returns a lazy pager. Keep the client in a named variable until iteration finishes; constructing it inline can let it be closed before the pager sends its request. The endpoint can still list deprecated preview image IDs, so cross-check the official deprecation page and add only current IDs.
- Probe image settings on the exact model and API surface. For `gemini-3.1-flash-image`, the minimum `generateContent` value is `ImageConfigDict(image_size='512')`; the superficially similar literal `'0.5K'` is invalid and returns HTTP 400. `gemini-3.1-flash-lite-image` supports only 1K output. Do not transfer value spellings between model families or API examples without a live check.

### Others

Not yet documented here. **When you add the next model for one of these providers, add the landmines you encountered to this section before closing the session** (see Step 9).

## Step 9 — Update this skill

After completing the model-add, before closing the session: if anything came up that isn't already documented in this skill — a new test that ratcheted, a provider-specific dispatch tuple, a misleading SDK behavior, a corrected misconception, an iteration the user had to walk you through — **add it to this SKILL.md**.

Specifically:
- Provider-specific landmines → the matching subsection (or create it).
- Generic process gaps → the relevant numbered step.
- Workflow shape errors → restructure the steps.

This skill exists to compound learnings. A model-add that surfaced new friction and didn't update this file wasted that friction.
