Agent skill

Ax AI

by dosco in dosco/aithy

This skill helps an LLM generate correct AI provider setup and configuration code using @ax-llm/ax.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Ax AI

skills CLI
$ npx skills add dosco/aithy --skill ax-ai -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dosco/aithy ax-ai --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ax-ai .claude/skills/ax-ai && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ax-ai
GitHub stars
107
Token cost
~8.2k tokens
SKILL.md length
2,850 words
Files
1
Skills in repo
17
Repo updated
First seen
Licence
Apache-2.0

At a glance

This skill helps an LLM generate correct AI provider setup and configuration code using @ax-llm/ax.

  • The user asks about ai()
  • SKILL.md covers Quick Setup, Renewable Credentials, Model Presets and Model Catalog, plus 17 more sections
  • Needs GOOGLE_VERTEX_ACCESS_TOKEN and PROVIDER_API_KEY
  • Adaptive balancing

What it does

Ax AI is an agent skill from dosco/aithy. This skill helps an LLM generate correct AI provider setup and configuration code using @ax-llm/ax. Use when the user asks about ai(), providers, models, routing, adaptive balancing, presets, embeddings, batch audio with ai.transcribe() or ai.speak(), extended thinking, context caching, or mentions OpenAI/Anthropic/Google/Azure/DeepSeek/Mistral/Cohere/Reka/Grok with @ax-llm/ax.

Its SKILL.md is about 8.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Embeddings, Transcription and Caching. It works with DeepSeek, Mistral AI, OpenAI and Microsoft Azure. The repository describes itself as: A personal AI agent that can work safely on your machine, remember useful context, and keep its data under your control. The licence is Apache-2.0.

When your agent uses it

  • The user asks about ai()
  • Adaptive balancing
  • Batch audio with ai.transcribe()
  • Extended thinking

Example prompts

  • “/ax-ai”

Requirements

  • A credential in PROVIDER_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 0c9855f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • raw.githubusercontent.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_VERTEX_ACCESS_TOKEN
    • PROVIDER_API_KEY
    • BEDROCK_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ax AI loads about 8.2k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 2,850 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~8.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dosco/aithy at commit 0c9855f, republished under its Apache-2.0 licence (© dosco). 2,850 words, ~8,180 tokens.

Download SKILL.mdSave it as .claude/skills/ax-ai/SKILL.md (or your agent's skills folder).
name
ax-ai
description
This skill helps an LLM generate correct AI provider setup and configuration code using @ax-llm/ax. Use when the user asks about ai(), providers, models, routing, adaptive balancing, presets, embeddings, batch audio with ai.transcribe() or ai.speak(), extended thinking, context caching, or mentions OpenAI/Anthropic/Google/Azure/DeepSeek/Mistral/Cohere/Reka/Grok with @ax-llm/ax.
version
24.0.16

AI Provider Codegen Rules (@ax-llm/ax)

Use this skill to generate AI provider setup, configuration, and chat code. Prefer short, modern, copyable patterns. Do not write tutorial prose unless the user explicitly asks for explanation.

Quick Setup

typescript
import { ai } from '@ax-llm/ax';

const openai = ai({ name: 'openai', apiKey: 'sk-...' });
const claude = ai({ name: 'anthropic', apiKey: 'sk-ant-...' });
const gemini = ai({ name: 'google-gemini', apiKey: 'AIza...' });
const azure = ai({ name: 'azure-openai', apiKey: 'your-key', resourceName: 'your-resource', deploymentName: 'gpt-5-4-mini' });
const deepseek = ai({ name: 'deepseek', apiKey: 'sk-...' });
const mistral = ai({ name: 'mistral', apiKey: 'your-key' });
const cohere = ai({ name: 'cohere', apiKey: 'your-key' });
const custom = ai({
  name: 'openai-compatible',
  apiKey: process.env.PROVIDER_API_KEY,
  apiURL: 'https://example.com/v1',
  config: { model: 'provider/model-name' },
});
const reka = ai({ name: 'reka', apiKey: 'your-key' });
const grok = ai({ name: 'grok', apiKey: 'your-key' });

name selects a deployment profile; config.model selects a model inside that deployment. Never infer provider behavior from a model ID. For example, name: 'together' with a deepseek-ai/... model uses Together's profile rules, not DeepSeek's native request shape.

Use axAIProfiles() to discover the complete named catalog and axGetAIProfile(name) to inspect endpoint requirements, authentication, operations, capabilities, model rules, sources, and review dates. Use name: 'openai-compatible' plus apiURL for an unlisted custom endpoint; unknown names are errors.

Profile-only branded classes were removed in the major-version migration. Use ai({ name: ... }) for Azure OpenAI, Cohere, DeepSeek, DeepSeek Responses, Mistral, Reka, and Grok. Retained low-level classes represent genuine transports/runtimes only; legacy model enum and catalog exports remain usable.

Renewable Credentials

Use credentialProvider for expiring bearer tokens. The callback runs for each request attempt and receives { profile, operation, method, url }. Fresh headers override static authentication headers.

typescript
const vertex = ai({
  name: 'vertex-ai',
  apiURL: process.env.VERTEX_AI_API_URL!,
  config: { model: 'google/gemma-4-26b-a4b-it-maas' },
  credentialProvider: async ({ operation, url }) => ({
    Authorization: `Bearer ${await tokenSource.fresh({ operation, url })}`,
  }),
});
  • Authentication-required profiles accept apiKey or credentialProvider.
  • The hook covers chat, stream, embeddings, Responses, transcription, speech, and retry attempts. A hook error stops before transport.
  • Do not expect an automatic replay after a completed 401 or 403; generation requests are not safely idempotent.
  • Keep ADC or cloud-SDK dependencies in the host. Ax core only owns the neutral callback contract.

Vertex capability rules are model-aware. Documented Gemini MaaS IDs prefer native schema. The exact google/gemma-4-26b-a4b-it-maas rule prefers json_object, excludes native schema, defaults thinking to max, writes chat_template_kwargs.enable_thinking, and extracts/replays reasoning_content. Unknown Vertex models remain conservative.

<!-- axir-nonportable:start webllm -->

WebLLM is browser-only and requires a host-created WebLLM engine. The host loads or reloads models with WebLLM APIs such as CreateMLCEngine(...); Ax only forwards chat requests to that loaded engine. Do not present WebLLM as a portable AxIR provider or a server-side default.

typescript
import { ai, AxAIWebLLMModel } from '@ax-llm/ax';

const engine = await CreateMLCEngine(AxAIWebLLMModel.Llama32_3B_Instruct);
const llm = ai({
  name: 'webllm',
  engine,
  config: {
    model: AxAIWebLLMModel.Llama32_3B_Instruct,
    stream: false,
    supportsFunctions: false,
  },
});
<!-- axir-nonportable:end webllm -->

Model Presets

typescript
import { ai, AxAIGoogleGeminiModel } from '@ax-llm/ax';

const gemini = ai({
  name: 'google-gemini',
  apiKey: process.env.GOOGLE_APIKEY!,
  config: { model: 'simple' },
  models: [
    { key: 'tiny', model: AxAIGoogleGeminiModel.Gemini35FlashLite, description: 'Fast + cheap', config: { maxTokens: 1024 } },
    { key: 'simple', model: AxAIGoogleGeminiModel.Gemini37Flash, description: 'Balanced' },
  ],
});

await gemini.chat({ model: 'tiny', chatPrompt: [{ role: 'user', content: 'Hi' }] });

Model Catalog

typescript
import { axGetSupportedAIModels } from '@ax-llm/ax';

const providers = axGetSupportedAIModels();
const openai = providers.find((provider) => provider.name === 'openai');
console.log(openai?.models[0]?.promptTokenCostPer1M);
console.log(openai?.capabilities.serviceTiers);

const reasoningModel = openai?.models.find(
  (model) => model.capabilities.thinkingLevels.includes('high')
);
console.log(reasoningModel?.capabilities.thinkingLevels);

const textProviders = axGetSupportedAIModels({ type: 'text' });
const embeddingProviders = axGetSupportedAIModels({ type: 'embeddings' });

Use axGetSupportedAIModels() to build provider/model selectors before creating an ai(...) instance. It returns bundled static metadata: provider names, display names, default models, raw AxModelInfo pricing/details, model type ('text', 'embeddings', 'code', or 'audio'), and normalized capabilities for thinking, thoughts, structured outputs, audio, temperature, top-p, portable thinking levels, and verified service tiers. Provider capabilities describe the default deployment profile; each static model also carries its resolved capabilities.thinkingLevels and capabilities.serviceTiers. Portable thinking levels can collapse onto the same provider-native value. serviceTiers lists verified explicit tiers; serviceTier: 'auto' remains available as the provider-delegated policy.

Provider groups and models are sorted cheapest to most expensive based on bundled input + output token pricing; unpriced models sort last. Dynamic profiles remain useful even when models is empty because their provider-level capability metadata is still returned.

Filter with { type: 'all' | 'text' | 'embeddings' | 'code' | 'audio' } or an array of those values. The 'text' filter includes code-capable models; use 'code' to show only code-first models.

Dynamic providers such as Azure OpenAI deployments are marked with isDynamic: true and may have an empty or static-limited model list.

Portable Inference Service Tiers

Use the shared serviceTier option with auto, standard, flex, or priority. Set it per call, as an instance option, or on a model-key preset; the narrower setting wins. auto delegates to the provider and is omitted when that provider has no explicit auto value. Ax rejects unsupported explicit tiers before transport.

typescript
import { ai, AxAIGoogleGeminiModel } from '@ax-llm/ax';

const gemini = ai({
  name: 'google-gemini',
  apiKey: process.env.GOOGLE_APIKEY!,
  config: { model: AxAIGoogleGeminiModel.Gemini35Flash },
});

const response = await gemini.chat(
  { chatPrompt: [{ role: 'user', content: 'Run this when capacity is free.' }] },
  { serviceTier: 'flex' },
);

Inspect ai.getFeatures(model).serviceTiers for the verified tiers on a named profile or exact model. Named OpenAI, Gemini, Azure, Bedrock, Mistral, Groq, OpenRouter, DeepInfra, Cerebras, Grok, Databricks, and Fireworks profiles map the portable value to their provider dialect. Exact modelInfo.supported metadata can opt a conservative custom deployment into a tier.

The applied provider value is normalized into response.modelUsage.tokens.serviceTier; aliases such as default, on_demand, and performance become standard or priority. Add serviceTierPricing.flex or .priority to AxModelInfo when tier pricing differs, and cost estimates will use the tier that actually served the request. Gemini service tiers remain GenerateContent-only: Vertex AI and Gemini Live reject explicit tiers. Anthropic speed: 'fast' remains a separate API.

Routing And Balancing

Choose the primitive by responsibility:

  • AxMultiServiceRouter combines model lists and dispatches the model key the caller already selected. It does not select a model.
  • AxProviderRouter selects a provider by request capability and may degrade unsupported media through configured processors. When the selected provider supports images natively, each image remains an image object with its payload, MIME type, detail level, cache and optimization hints, alt text, and ordering with surrounding text intact.
  • AxBalancer without a strategy orders equivalent services once with a comparator, retries transient provider failures, and fails over in that order.
  • AxBalancer with strategy.type: 'adaptive' selects among services exposing the same logical model aliases using learned provider reliability, successful latency, and estimated cost.

Adaptive balancing is operational routing, not semantic prompt-to-model routing. Every provider model behind an alias must be an acceptable substitute for that application. Keep quality evaluation and content-aware model selection outside the balancer.

typescript
import { AxBalancer, AxInMemoryBalancerStatsStore } from '@ax-llm/ax';

const statsStore = new AxInMemoryBalancerStatsStore();
const routeKeys = new Map<string, string>([
  [openai.getId(), 'openai-primary'],
  [anthropic.getId(), 'anthropic-primary'],
]);

const llm = AxBalancer.create([openai, anthropic] as const, {
  strategy: {
    type: 'adaptive',
    deadlineMs: 6_000,
    badOutcomeCost: 0.02,
    expectedTokens: { promptTokens: 1_200, completionTokens: 300 },
    namespace: 'support-v1',
    routeKey: (service) => {
      const key = routeKeys.get(service.getId());
      if (!key) throw new Error('Missing stable route key.');
      return key;
    },
    slice: ({ options }) =>
      options?.customLabels?.workflow ?? 'default-workflow',
    statsStore,
    onRoutingEvent: (event) => telemetry.emit('llm.route', event),
  },
});

The score is estimated request cost plus badOutcomeCost times the probability of provider failure or missing deadlineMs. badOutcomeCost and estimated cost must use the same currency or unit. By default, cost uses expectedTokens, the route's concrete model mapping, and getEstimatedCost(); missing catalog pricing contributes zero, while estimateCost can supply application pricing. Failures use an EWMA; successful latency is modeled in log space with a Normal-Inverse-Gamma posterior, and Thompson sampling supplies the deadline risk. Capability filtering still runs before ranking.

Rules:

  • Reuse the in-memory store across balancers in one process. For multiple processes, implement AxBalancerStatsStore with Redis or an application database; its observe() operation must be atomic.
  • A custom store requires stable, unique routeKey values. Stats are partitioned by namespace, slice, logical model, and route.
  • statsStore is decision state. onRoutingEvent is best-effort telemetry and must not be used as the authoritative routing state.
  • Routing events contain scores, route metadata, and sanitized failure categories, never prompts, responses, or raw provider errors.
  • Adaptive mode attempts each candidate once. Provider-client retries remain controlled by AxAIServiceOptions.retry.
  • Streaming can fail over before the first emitted chunk. A mid-stream transient failure is learned but cannot be replayed after partial output; caller cancellation is not recorded.
  • Adaptive selection applies to chat. Embedding, transcription, and speech keep existing balancer behavior.

See the adaptive balancer example for complete provider setup.

Chat

typescript
const res = await llm.chat({
  chatPrompt: [
    { role: 'system', content: 'You are concise.' },
    { role: 'user', content: 'Write a haiku about the ocean.' },
  ],
});
console.log(res.results[0]?.content);

Batch Audio

Use ai.transcribe(...) for batch speech-to-text and ai.speak(...) for batch text-to-speech. These are separate from conversational .chat() audio config.

typescript
const transcript = await llm.transcribe({
  audio: { data: base64Wav, format: 'wav' },
  model: 'gpt-4o-mini-transcribe',
  language: 'en',
});

const speech = await llm.speak({
  text: transcript.text,
  model: 'gpt-4o-mini-tts',
  voice: 'alloy',
  format: 'mp3',
});

console.log(transcript.text);
console.log(speech.data);

Providers without the requested audio endpoint throw AxMediaNotSupportedError. Use speech forward options for signature audio artifacts and modelConfig.audio for conversational chat audio.

Common Options

  • stream (boolean): enable SSE; true by default
  • thinkingTokenBudget: 'minimal' | 'low' | 'medium' | 'high' | 'highest' | 'none'
  • showThoughts: include thoughts in output
  • functionCallMode: 'auto' | 'native' | 'prompt'
  • debug, logger, tracer, rateLimiter, timeout

Global Runtime Defaults

Use axGlobals when the app wants one live default for AI requests, generator runs, flows, or metrics:

typescript
import { ai, axGlobals, axCreateDefaultColorLogger } from '@ax-llm/ax';
import { metrics, trace } from '@opentelemetry/api';

axGlobals.rateLimiter = async (next, info) => next();
axGlobals.tracer = trace.getTracer('my-app');
axGlobals.meter = metrics.getMeter('my-app');
axGlobals.debug = true;
axGlobals.logger = axCreateDefaultColorLogger();
axGlobals.customLabels = { service: 'api' };
axGlobals.onUsage = (event) => usageQueue.enqueue(event);

const llm = ai({ name: 'openai', apiKey: process.env.OPENAI_APIKEY! });

Rules:

  • axGlobals.rateLimiter, tracer, meter, logger, debug, abortSignal, and customLabels are live runtime defaults; each operation snapshots them at its start, even if the AI instance already exists.
  • Precedence is: per-call options, then explicit AI/service options, then current axGlobals, then built-in defaults.
  • The limiter receives next plus operation, provider, model, streaming state, and previous service usage. It wraps chat and embedding provider execution, including streaming and retries; its errors propagate. It may delay, reject, skip, or invoke next multiple times.
  • Tracer, meter, and usage-observer failures are fail-open. Limiter failures are fail-closed.
  • Runtime-hook telemetry contains metadata and usage only, never prompts, outputs, tool arguments, or tool results.
  • External meter instruments are independent of balancer-local getMetrics() snapshots. Adapt AxMeter to OpenTelemetry at the application boundary; generated packages do not require an OpenTelemetry dependency.
  • customLabels merge from globals to service to call options; later sources override earlier keys.
  • abortSignal values are merged, so either a global shutdown signal or a local request signal can cancel the request.
  • axGlobals.onUsage receives one immutable normalized event for each completed chat or embedding call that reports token usage. A fully consumed stream emits once.
  • Usage observers are best-effort and fail-open. Ax does not await them; synchronously enqueue events and persist or aggregate them out of band.

Clear process-wide hooks during shutdown or test teardown:

typescript
axGlobals.rateLimiter = undefined;
axGlobals.tracer = undefined;
axGlobals.meter = undefined;

Use usageContext for multi-tenant and request attribution:

typescript
const llm = ai({
  name: 'openai',
  apiKey: process.env.OPENAI_APIKEY!,
  options: {
    usageContext: {
      tenantId: 'tenant-42',
      feature: 'support-chat',
      attributes: { environment: 'production' },
    },
  },
});

await llm.chat(request, {
  usageContext: {
    userId: user.id,
    requestId: requestId,
    runId: runId,
  },
});

Per-call context overrides service defaults, while attributes are shallow-merged. Events include normalized tokens, provider/model, available session and remote IDs, and a streaming flag. They do not estimate currency cost; calculate that downstream against a versioned pricing table.

DeepSeek Notes

typescript
import { ai, AxAIDeepSeekModel } from '@ax-llm/ax';

const deepseek = ai({
  name: 'deepseek',
  apiKey: process.env.DEEPSEEK_APIKEY!,
  config: { model: AxAIDeepSeekModel.DeepSeekV4Flash },
});

DeepSeek's current API models are deepseek-v4-flash and deepseek-v4-pro. Legacy model enum/catalog values remain available for source compatibility, but profile rules are applied only to model IDs verified for the selected deployment.

DeepSeek V4 supports thinking mode. When thinkingTokenBudget is omitted, Ax selects its logical max level and sends thinking: { type: "enabled" } with reasoning_effort: "max". Set thinkingTokenBudget: "none" explicitly to disable it. DeepSeek's API exposes low, medium, high, and max: Ax maps minimal and low to low, preserves medium, maps high to high, and maps highest to max. DeepSeek V4 thinking models support tools, but reject the tool_choice request parameter, so Ax omits auto and Ax-generated __axOutput tool choices for deepseek-v4-pro, deepseek-v4-flash, and deepseek-reasoner while still sending tool definitions. An explicitly forced caller tool choice fails before the request. DeepSeek does not support native JSON schema structured outputs. Ax therefore uses validated json_object for a single required string or code output, including an AxAgent actor's javascriptCode field, without exposing a provider tool. Richer structured outputs use the synthetic __axOutput function when function calling is available, or validated json_object when it is not.

DeepSeek Chat returns thinking traces as reasoning_content. During a tool loop, Ax preserves that field on the assistant tool-call message and sends it back on the following request together with non-null content and the original tool calls. This compatibility mode is declared by the DeepSeek deployment profile; the official OpenAI Chat adapter does not emit or expose reasoning_content.

The same logical default is declared independently for verified DeepSeek V4 rules in the Together, Fireworks, and OpenRouter profiles, then mapped to each deployment's own request dialect. A custom openai-compatible endpoint never inherits it from a DeepSeek-looking model ID.

Other verified deployment rules follow the same policy. Grok 4.6 maps logical max to xhigh, while Grok 4.5 and 4.3 map it to high. Groq GPT-OSS and Cerebras GPT-OSS map it to high; Groq Qwen 3.6 maps it to its documented reasoning-enabled default; Cerebras Gemma 4 and DeepInfra DeepSeek R1 map it to high. An explicit none is sent only for model/deployment combinations that document disabling reasoning. Grok 4.6/4.5 and GPT-OSS on Groq or Cerebras reject none before network I/O because those APIs do not support disabling reasoning for those models.

Hugging Face Router remains conservative: routing policies such as :fastest may choose a different inference provider without changing the base model ID, so Ax does not attach one provider's reasoning contract to that dynamic route.

Show full SKILL.md (1,066 more words)Show less

Extended Thinking

typescript
import { ai, AxAIAnthropicModel } from '@ax-llm/ax';

const claude = ai({
  name: 'anthropic',
  apiKey: process.env.ANTHROPIC_APIKEY!,
  config: { model: AxAIAnthropicModel.Claude48Opus },
});

const res = await claude.chat(
  { chatPrompt: [{ role: 'user', content: 'Solve step by step...' }] },
  { thinkingTokenBudget: 'medium', showThoughts: true },
);
console.log(res.results[0]?.thought);
console.log(res.results[0]?.content);
Budget Levels
LevelAnthropic (tokens)Gemini 2.5 (tokens)Gemini 3 level
'none'disabled0 on Flash/Lite; minimum on Prolowest supported, thoughts hidden
'minimal'1,024200minimal, or low when minimal is unsupported
'low'5,000800low
'medium'10,0005,000medium, or the nearest image/legacy level
'high'20,00010,000high
'highest'32,00024,500high

Gemini 3 uses thinkingLevel; Gemini 2.5 and older models use numeric thinkingBudget. Ax selects the wire field after resolving a named model preset to its real model. Gemini 3.7 Flash and Gemini 3.1 Pro clamp minimal to low; image and legacy Gemini 3 models clamp to their documented two-level sets. Numeric Gemini 3 budgets fail locally. none always hides returned thoughts, even when the model must still perform its minimum amount of thinking.

The native google-gemini deployment profile and its aliases use these Gemini rules, including when configured for Vertex with projectId and region. The separate OpenAI-compatible vertex-ai profile keeps its own request rules and does not inherit native Gemini fields from a Gemini-looking model ID.

For GPT-5.6, these map to none, low, low, medium, high, and a top rung that depends on the API surface: xhigh on Chat Completions, which rejects max, and max on the Responses API, which is the only place it is served. Earlier OpenAI models retain their existing mapping.

Anthropic Model-Specific Behavior
  • Opus 4.8, 4.7, and 4.6 plus Sonnet 5: adaptive thinking, no manual budget_tokens, and no temperature / topP / topK. When thoughts are requested, Ax asks Anthropic for summarized display; when they are hidden, Ax explicitly requests display: 'omitted'.
  • Opus 4.5: budget_tokens + effort levels (capped at 'high')
  • Other thinking models: budget tokens only

Anthropic modelConfig.effort can be set directly on a request. Fast mode and task budgets are Anthropic-only opt-ins; taskBudget.total must be at least 20,000 tokens.

typescript
const res = await claude.chat({
  chatPrompt: [{ role: 'user', content: 'Review this migration plan.' }],
  modelConfig: {
    effort: 'xhigh',
    speed: 'fast',
    taskBudget: { type: 'tokens', total: 64_000 },
  },
});
Custom Thinking Levels
typescript
const claude = ai({
  name: 'anthropic',
  apiKey: '...',
  config: {
    model: AxAIAnthropicModel.Claude48Opus,
    thinkingTokenBudgetLevels: {
      minimal: 2048,
      low: 8000,
      medium: 16000,
      high: 25000,
      highest: 40000,
    },
    effortLevelMapping: {
      minimal: 'low',
      low: 'medium',
      medium: 'high',
      high: 'high',
      highest: 'max',
    },
  },
});

Embeddings

typescript
const { embeddings } = await llm.embed({
  texts: ['hello', 'world'],
  embedModel: 'text-embedding-005',
});

Vertex AI Locations

When projectId and region are set for Google Gemini or Anthropic on Vertex AI, Ax selects the service hostname from the location automatically:

  • global uses aiplatform.googleapis.com
  • us and eu use the multi-region .rep.googleapis.com endpoints
  • regional locations such as us-central1 use {region}-aiplatform.googleapis.com

Pass the canonical lower-case Vertex location ID. Ax preserves the supplied value and does not normalize or validate it.

The generated Python, Java, C++, Go, and Rust clients accept the same projectId / project_id, region, and optional endpointId / endpoint_id options. In generated clients, apiKey / api_key is a caller-supplied bearer access token (or GOOGLE_VERTEX_ACCESS_TOKEN); ADC discovery and automatic token refresh remain host-owned. An explicit baseUrl / base_url always wins.

Context Caching

typescript
const result = await gen.forward(llm, { code, language }, {
  mem,
  sessionId: 'code-review-session',
  contextCache: {
    ttlSeconds: 3600,
    cacheBreakpoint: 'after-examples',
  },
});

Breakpoint values: 'system' | 'after-functions' | 'after-examples'

Provider behavior:

  • Google Gemini: explicit caching with cache resource ID and auto TTL refresh; failed refreshes recreate or fall back uncached, and rejected Ax-managed caches retry once without the cache
  • Anthropic: implicit via cache_control markers
  • OpenAI: explicit prompt_cache_breakpoint markers, GPT-5.6+ only. Earlier families cache automatically and predate the parameters, so nothing is sent to them. Only the openai provider opts in — Azure OpenAI shares the request builder and the same model enum, so a gpt-5.6-* deployment sends nothing, and openai-responses does not send breakpoints either (it does report cacheCreationTokens, which is provider-wide)
OpenAI prompt cache keys

GPT-5.6+ needs a key that is stable per conversation to match reliably; it routes the request to the shard the cache lives on. Set promptCacheKey, or let it fall back to sessionId. Keep it under roughly 15 requests/minute per key.

typescript
const result = await gen.forward(llm, values, {
  mem,
  promptCacheKey: `review:${pullRequestId}`,
  contextCache: {},
});

AxGen forwards these provider options after merging program defaults with the per-call options. Generated language packages preserve the same promptCacheKey / sessionId / contextCache forwarding contract.

Markers must not move. A breakpoint marker is part of its content block, so marking only "the newest stable message" each turn un-marks what the previous turn marked, changing the prefix and voiding the entry that turn wrote. Ax marks by absolute index from the front, which is stable for an append-only conversation. Anything that rewrites the front of the history — dynamically added functions changing the system prompt, or mem.rewindToTag — costs a cache miss.

Keep caching on for the whole conversation. The provider marks all or nothing, so a turn that omits contextCache sends the prompt unmarked and the next turn rewrites the cache from scratch.

External Registry (serverless)
typescript
const accountId = getRequiredAccountId();
const registry: AxContextCacheRegistry = {
  get: async (key) => {
    const value = await redis.get(`context-cache:${accountId}:${key}`);
    return value ? JSON.parse(value) : undefined;
  },
  set: async (key, entry) => {
    const ttl = Math.max(1, Math.ceil((entry.expiresAt - Date.now()) / 1000));
    await redis.set(
      `context-cache:${accountId}:${key}`,
      JSON.stringify(entry),
      { ex: ttl }
    );
  },
};

Ax registry keys are content-based and are not account-scoped. Require a stable tenant/account namespace when cross-account cache sharing is unsafe; do not silently fall back to a global namespace.

AWS Bedrock

Use the amazon-bedrock profile for Bedrock's OpenAI-compatible Mantle endpoint. Supply the account/region-specific OpenAI base URL and a Bedrock API key explicitly:

typescript
const bedrock = ai({
  name: 'amazon-bedrock',
  apiURL: process.env.BEDROCK_OPENAI_BASE_URL!,
  apiKey: process.env.BEDROCK_API_KEY!,
  config: { model: process.env.BEDROCK_MODEL_ID! },
});

The separate AWS package remains available when native AWS SDK authentication, regional fallback, or non-Mantle Bedrock behavior is required:

typescript
import { AxAIBedrock, AxAIBedrockModel } from '@ax-llm/ax-ai-aws-bedrock';

const bedrock = new AxAIBedrock({
  region: 'us-east-2',
  fallbackRegions: ['us-west-2'],
  config: { model: AxAIBedrockModel.ClaudeOpus45 },
});

Vercel AI SDK Integration

typescript
import { generateText } from 'ai';
import { ai } from '@ax-llm/ax';
import { AxAIProvider } from '@ax-llm/ax-ai-sdk-provider';

const axAI = ai({
  name: 'openai',
  apiKey: process.env.OPENAI_APIKEY ?? '',
});
const model = new AxAIProvider(axAI);

const result = await generateText({
  model,
  prompt: 'Hello!',
});

MCP + AxJSRuntime

typescript
import { AxMCPClient } from '@ax-llm/ax';
import { axCreateMCPStdioTransport } from '@ax-llm/ax-tools';

const transport = axCreateMCPStdioTransport({
  command: 'npx',
  args: ['-y', '@anthropic/mcp-server-filesystem'],
});
const client = new AxMCPClient(transport);

For server notifications, call client.startListening({ signal, onError }) or attach the client through AxMCPEventSource. The event adapter is preferred for autonomous work because protocol callbacks only enqueue; explicit routes decide whether to observe, invalidate, resume, or wake.

For signed UCP lifecycle requests, mount AxUCPWebhookEventSource.ingest(request) in application-owned HTTP hosting. Signature, profile, digest, freshness, and replay verification completes before the event runtime sees the request.

Critical Rules

  • Use ai() factory for all providers.
  • Use axAIProfiles() as the source of truth for names. Core names include 'openai', 'openai-compatible', 'openai-responses', 'anthropic', 'google-gemini', 'azure-openai', 'deepseek', 'mistral', 'cohere', 'grok', routers such as 'together', 'openrouter', and 'orcarouter', hosted inference profiles, and configurable local runtimes.
  • Thinking constraints on Anthropic: every adaptive-thinking model omits temperature, topP, and topK; older thinking models ignore temperature and topK, with topP only sent if >= 0.95.
  • ai({ name: 'amazon-bedrock', apiURL: ... }) targets Bedrock's OpenAI-compatible endpoint. new AxAIBedrock() remains the separate AWS-native runtime client.
  • Vercel AI SDK uses AxAIProvider wrapper.

Examples

Fetch these for full working code:

Do Not Generate

  • Do not use new AxAIOpenAI(...) or similar class constructors for standard providers; use ai().
  • Do not hardcode provider class names when ai({ name: ... }) covers the provider.
  • Do not mix thinkingTokenBudget with explicit temperature on Anthropic thinking models.
  • Do not confuse the amazon-bedrock OpenAI-compatible profile with the separate AWS-native AxAIBedrock client.
  • Do not omit resourceName and deploymentName for Azure OpenAI.

© dosco, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/ax-ai of dosco/aithy.

Open the folder on GitHubat commit 0c9855f

Compare with similar skills

Ax AI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ax AI compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ax AI this skilldosco/aithy107—~8.2kAutomated safety check: PassApache-2.0
AI SDK Developmenttrypostit/trypost6851 repos~3.5kAutomated safety check: PassMIT
Keiroutermydisha/keirouter147—~995Automated safety check: PassMIT
Azure AI Openai Dotnetmicrosoft/skills3.1k5 repos~3.4kAutomated safety check: PassMIT
Embeddings via 9Routerdecolua/9router30k—~604Automated safety check: PassMIT
Caching Architecturemajiayu000/litellm-rs117—~2kAutomated safety check: PassMIT

Similar skills

  • AI SDK Development

    trypostit/trypost

    TRIGGER when working with ai-sdk which is Laravel official first-party AI SDK.

    685 GitHub starsUsed in 1 repo~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Keirouter

    mydisha/keirouter

    Entry point for KeiRouter — local/remote AI gateway with OpenAI-compatible REST for chat, image, TTS, embeddings, web search, web fetch.

    147 GitHub stars~995 tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check passed
  • Azure AI Openai Dotnet

    microsoft/skills

    Official

    Azure OpenAI SDK for .NET. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 5 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    30k GitHub stars~604 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Caching Architecture

    majiayu000/litellm-rs

    LiteLLM-RS response caching architecture. An agent skill from majiayu000/litellm-rs.

    117 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Awesome Free LLM APIs

    LeoYeAI/openclaw-master-skills

    Reference guide for permanent free-tier LLM APIs with rate limits, model lists, and OpenAI-compatible integration patterns.

    2.2k GitHub stars~4.1k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from dosco/aithy

All 17 skills in this repo
  • This skill helps an LLM generate correct AxAgent observability code using @ax-llm/ax.

    107 GitHub stars~4.4k tokensUpdated 1 mo ago
    Auto-check passed
  • This skill helps an LLM generate correct AxAgent tuning and evaluation code using @ax-llm/ax.

    107 GitHub stars~4.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Ax Audio

    dosco/aithy

    This skill helps an LLM generate correct audio code with @ax-llm/ax.

    107 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Ax Gepa

    dosco/aithy

    This skill helps an LLM generate correct AxGEPA optimization code using @ax-llm/ax.

    107 GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Ax LLM

    dosco/aithy

    This skill helps with using the @ax-llm/ax TypeScript library for building LLM applications.

    107 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Ax MCP

    dosco/aithy

    This skill helps an LLM build correct native Model Context Protocol integrations with @ax-llm/ax.

    107 GitHub stars~4.4k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Ax AI

What does Ax AI do?

This skill helps an LLM generate correct AI provider setup and configuration code using @ax-llm/ax. Ax AI is an agent skill from dosco/aithy. This skill helps an LLM generate correct AI provider setup and configuration code using @ax-llm/ax.

When should I use Ax AI?

Ax AI fits situations like: the user asks about ai(); adaptive balancing; batch audio with ai.transcribe(); extended thinking.

How do I install Ax AI in Claude Code?

Run `npx skills add dosco/aithy --skill ax-ai -a claude-code`. Or copy the skill folder (.claude/skills/ax-ai in dosco/aithy) into .claude/skills/ax-ai in your project. Claude Code loads it when a task matches its description.

How do I install Ax AI in Codex?

Run `npx skills add dosco/aithy --skill ax-ai -a codex`. Or copy the skill folder (.claude/skills/ax-ai in dosco/aithy) into .agents/skills/ax-ai in your project. Codex loads it when a task matches its description.

Can I use Ax AI in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dosco/aithy --skill ax-ai -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ax-ai, .gemini/skills/ax-ai, .github/skills/ax-ai and .opencode/skills/ax-ai in your project.

What does Ax AI need to run?

Going by SKILL.md and its folder, Ax AI needs credentials named GOOGLE_VERTEX_ACCESS_TOKEN, PROVIDER_API_KEY and BEDROCK_API_KEY. Our summary lists: A credential in PROVIDER_API_KEY.

Does Ax AI access the network?

SKILL.md names 1 domain. As links in the text: raw.githubusercontent.com. This is read from the text; nothing was executed.

Is Ax AI safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ax AI use?

Ax AI is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ax AI use?

About 8.2k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ax AI?

Skills that share tags, products or a category with Ax AI: AI SDK Development (trypostit/trypost, 685 stars), Keirouter (mydisha/keirouter, 147 stars), Azure AI Openai Dotnet (microsoft/skills, 3.1k stars) and Embeddings via 9Router (decolua/9router, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ax AI?

dosco (a GitHub user) maintains it in dosco/aithy, which has 107 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on August 31, 2026.

Source: dosco/aithy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.