Agent skill

Ax Audio

by dosco in dosco/aithy

This skill helps an LLM generate correct audio code with @ax-llm/ax.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Ax Audio

skills CLI
$ npx skills add dosco/aithy --skill ax-audio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dosco/aithy ax-audio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ax-audio .claude/skills/ax-audio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ax-audio
GitHub stars
107
Token cost
~2.5k tokens
SKILL.md length
547 words
Files
1
Skills in repo
17
Repo updated
First seen
Licence
Apache-2.0

At a glance

This skill helps an LLM generate correct audio code with @ax-llm/ax.

  • The user asks about ai.transcribe()
  • SKILL.md covers Core Rules, Direct Batch APIs, Signature Audio Artifacts and Agent Audio Inputs, plus 7 more sections
  • Needs GROK_API_KEY
  • Signature audio inputs

What it does

Ax Audio is an agent skill from dosco/aithy. This skill helps an LLM generate correct audio code with @ax-llm/ax. Use when the user asks about ai.transcribe(), ai.speak(), signature audio inputs or outputs, agent audio behavior, .chat() conversational audio, OpenAI audio or realtime models, Gemini Live native audio, Grok Voice Agent models, voices, formats, transcripts, or how audio fits with structured outputs.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Transcription, Structured output and tool calling and Speech recognition and synthesis. It works with OpenAI. The repository describes itself as: A personal AI agent that can work safely on your machine, remember useful context, and keep its data under your control. The licence is Apache-2.0.

When your agent uses it

  • The user asks about ai.transcribe()
  • Signature audio inputs
  • Agent audio behavior
  • .chat() conversational audio

Example prompts

  • “/ax-audio”

Requirements

  • A credential in GROK_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 0c9855f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GROK_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ax Audio loads about 2.5k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 547 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dosco/aithy at commit 0c9855f, republished under its Apache-2.0 licence (© dosco). 547 words, ~2,506 tokens.

Download SKILL.mdSave it as .claude/skills/ax-audio/SKILL.md (or your agent's skills folder).
name
ax-audio
description
This skill helps an LLM generate correct audio code with @ax-llm/ax. Use when the user asks about ai.transcribe(), ai.speak(), signature audio inputs or outputs, agent audio behavior, .chat() conversational audio, OpenAI audio or realtime models, Gemini Live native audio, Grok Voice Agent models, voices, formats, transcripts, or how audio fits with structured outputs.
version
24.0.16

Audio I/O Codegen Rules (@ax-llm/ax)

Use this skill for audio in Ax. Pick the smallest audio surface that matches the job:

  • Use ai.transcribe(...) for batch speech-to-text.
  • Use ai.speak(...) for batch text-to-speech.
  • Use speech:audio signature outputs for structured programs that should return synthesized audio artifacts.
  • Use .chat() audio config for conversational or realtime audio turns.

Core Rules

  • Input :audio is an audio input value: { data, format?, mimeType?, sampleRate?, channels? }.
  • Output :audio is a scripted audio artifact. The model returns plain text for that field; Ax synthesizes it after structured output parsing.
  • Output audio JSON schema is model-facing string, not a binary object.
  • Agents transcribe input audio fields before planner/executor/responder stages by default, so agent stages see text instead of base64 audio.
  • Realtime and conversational audio still use .chat() and modelConfig.audio.
  • Batch signature audio artifacts use forward-time speech options, not modelConfig.audio.

Direct Batch APIs

typescript
import { ai } from '@ax-llm/ax';

const llm = ai({ name: 'openai', apiKey: process.env.OPENAI_APIKEY! });

const transcript = await llm.transcribe({
  audio: { data: base64Wav, format: 'wav' },
  model: 'gpt-4o-mini-transcribe',
  language: 'en',
  prompt: 'Product support call',
});

const speech = await llm.speak({
  text: transcript.text,
  model: 'gpt-4o-mini-tts',
  voice: 'alloy',
  format: 'mp3',
});

console.log(transcript.text);
console.log(speech.data);
console.log(speech.transcript);

Providers without the requested batch audio capability throw AxMediaNotSupportedError.

Signature Audio Artifacts

typescript
import { ai, ax } from '@ax-llm/ax';

const llm = ai({ name: 'openai', apiKey: process.env.OPENAI_APIKEY! });
const say = ax('question:string -> speech:audio, summary:string');

const result = await say.forward(
  llm,
  { question: 'Explain retries in one sentence.' },
  {
    speech: {
      speak: { voice: 'alloy', format: 'mp3' },
      fields: {
        speech: { voice: 'alloy' },
      },
    },
  }
);

console.log(result.summary);
console.log(result.speech.data);
console.log(result.speech.mimeType);
console.log(result.speech.transcript);

The model emits a text script for speech; Ax replaces it with AxChatAudioOutput after result selection. If the field already contains an audio artifact with { data } or { id }, Ax leaves it alone.

Agent Audio Inputs

typescript
import { agent, ai } from '@ax-llm/ax';

const llm = ai({ name: 'openai', apiKey: process.env.OPENAI_APIKEY! });

const voiceAgent = agent(
  'recording:audio, question:string -> speech:audio, summary:string',
  {
    agentIdentity: {
      name: 'Voice Assistant',
      description: 'Answers spoken requests with spoken and written output',
    },
    contextFields: [],
  }
);

const result = await voiceAgent.forward(
  llm,
  {
    recording: { data: base64Wav, format: 'wav' },
    question: 'What should I do next?',
  },
  {
    speech: {
      transcribe: { model: 'gpt-4o-mini-transcribe' },
      speak: { voice: 'alloy', format: 'mp3' },
    },
  }
);

console.log(result.summary);
console.log(result.speech.data);

The agent runtime transcribes recording first and passes the transcript through the internal agent stages. Use direct ax(...) or .chat() when you specifically want native audio understanding in the model call.

Conversational .chat() Audio

Use modelConfig.audio for conversational audio turns where audio is part of the chat response instead of a structured signature field.

typescript
const res = await llm.chat({
  chatPrompt: [{ role: 'user', content: 'Say hello out loud.' }],
  modelConfig: {
    audio: { output: { enabled: true, voice: 'alloy', format: 'wav' } },
  },
});

console.log(res.results[0]?.content);
console.log(res.results[0]?.audio?.data);
console.log(res.results[0]?.audio?.transcript);

Config Shape

typescript
type AxAudioFormat =
  | 'wav'
  | 'mp3'
  | 'flac'
  | 'opus'
  | 'aac'
  | 'pcm16'
  | 'pcm'
  | 'ogg'
  | 'raw'
  | 'mulaw'
  | 'ulaw'
  | 'alaw';

type AxSpeechConfig = {
  transcribe?: {
    model?: string;
    language?: string;
    prompt?: string;
  };
  speak?: {
    model?: string;
    voice?: string;
    format?: AxAudioFormat;
  };
  fields?: Record<
    string,
    {
      model?: string;
      voice?: string;
      format?: AxAudioFormat;
    }
  >;
};

OpenAI Defaults

Use axAIOpenAIAudioDefaultConfig() for OpenAI request-based audio chat:

  • model: gpt-audio-mini
  • output enabled
  • voice: alloy
  • output format: wav
  • transcript enabled
  • streaming disabled by default
  • audio input formats: wav, mp3
  • audio output formats: wav, mp3, flac, opus, aac, pcm16
typescript
import { ai, axAIOpenAIAudioDefaultConfig } from '@ax-llm/ax';

const openai = ai({
  name: 'openai',
  apiKey: process.env.OPENAI_APIKEY!,
  config: axAIOpenAIAudioDefaultConfig(),
});

const res = await openai.chat({
  chatPrompt: [
    {
      role: 'user',
      content: [
        { type: 'text', text: 'What is in this recording?' },
        { type: 'audio', data: base64Wav, format: 'wav' },
      ],
    },
  ],
});

console.log(res.results[0]?.content);
console.log(res.results[0]?.audio?.data);

Use axAIOpenAIRealtimeDefaultConfig() for OpenAI realtime speech-to-speech:

  • model: gpt-realtime-2
  • output enabled
  • voice: marin
  • output format: pcm16
  • input default: audio/pcm, mono, 24000 Hz
  • turn timeout: 30000
  • streaming disabled by default

Use axAIOpenAIRealtimeTranscriptionDefaultConfig() for realtime transcript deltas:

  • model: gpt-realtime-whisper
  • input default: audio/pcm, mono, 24000 Hz
  • output audio disabled; transcript text is returned on content

Realtime models use a one-turn WebSocket call under .chat(). In Node, pass a WebSocket constructor through request options:

typescript
import WebSocket from 'ws';
import { ai, axAIOpenAIRealtimeDefaultConfig } from '@ax-llm/ax';

const openai = ai({
  name: 'openai',
  apiKey: process.env.OPENAI_APIKEY!,
  config: axAIOpenAIRealtimeDefaultConfig(),
});

const stream = await openai.chat(
  {
    chatPrompt: [{ role: 'user', content: 'Say hello out loud.' }],
  },
  { stream: true, webSocket: WebSocket }
);

For follow-up turns, keep the assistant audio reference in history:

typescript
await openai.chat({
  chatPrompt: [
    { role: 'assistant', audio: { id: previousAudioId } },
    { role: 'user', content: 'Repeat that more slowly.' },
  ],
});
Show full SKILL.md (185 more words)Show less

Gemini Live Defaults

Use axAIGoogleGeminiLiveAudioDefaultConfig() for Gemini native audio:

  • model: gemini-2.5-flash-native-audio-preview-12-2025
  • output enabled
  • voice: Kore
  • output format: pcm16
  • output sample rate: 24000
  • input default: audio/pcm;rate=16000, mono
  • transcript enabled
  • turn timeout: 30000
  • streaming disabled by default
typescript
import { ai, axAIGoogleGeminiLiveAudioDefaultConfig } from '@ax-llm/ax';

const gemini = ai({
  name: 'google-gemini',
  apiKey: process.env.GOOGLE_APIKEY!,
  config: axAIGoogleGeminiLiveAudioDefaultConfig(),
});

const res = await gemini.chat({
  chatPrompt: [
    {
      role: 'user',
      content: [
        { type: 'text', text: 'Answer this spoken question.' },
        {
          type: 'audio',
          data: base64Pcm16,
          format: 'pcm16',
          sampleRate: 16000,
          channels: 1,
        },
      ],
    },
  ],
});

console.log(res.results[0]?.content);
console.log(res.results[0]?.audio?.data);

Gemini Live uses a one-turn WebSocket call under .chat(). It expects PCM input for native audio turns; use format: 'pcm16' or mimeType: 'audio/pcm;rate=16000'.

Grok Voice Defaults

Use axAIGrokVoiceDefaultConfig() for xAI Grok Voice Agent:

  • model: grok-voice-think-fast-1.0
  • output enabled
  • voice: eve
  • output format: pcm16
  • output sample rate: 24000
  • input default: audio/pcm, mono, 24000 Hz
  • transcript enabled
  • turn timeout: 30000
  • streaming disabled by default
typescript
import WebSocket from 'ws';
import { ai, axAIGrokVoiceDefaultConfig } from '@ax-llm/ax';

const grok = ai({
  name: 'grok',
  apiKey: process.env.GROK_API_KEY!,
  config: axAIGrokVoiceDefaultConfig(),
});

const res = await grok.chat(
  {
    chatPrompt: [{ role: 'user', content: 'Say hello out loud.' }],
  },
  { webSocket: WebSocket }
);

console.log(res.results[0]?.content);
console.log(res.results[0]?.audio?.data);

Grok Voice uses a one-turn WebSocket call under .chat(). It expects PCM input for spoken input turns; use format: 'pcm16' or mimeType: 'audio/pcm'.

Streaming Audio

OpenAI audio chat, OpenAI Realtime, Gemini Live, and Grok Voice all default to non-streaming, but each can stream deltas when you pass { stream: true }.

typescript
const stream = await llm.chat(
  {
    chatPrompt: [{ role: 'user', content: 'Say hello.' }],
  },
  { stream: true }
);

for await (const chunk of stream) {
  const audio = chunk.results[0]?.audio;
  if (audio?.isDelta) {
    playAudioChunk(audio.data);
  }
}

Structured Outputs

Use signature audio outputs for structured speech artifacts:

typescript
const gen = ax('question:string -> answer:string, speech:audio');

Use .chat() audio when the response itself is a conversational audio turn. Do not combine .chat() audio output with provider-native structured response formats unless that provider explicitly supports the combination.

© dosco, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/ax-audio of dosco/aithy.

Open the folder on GitHubat commit 0c9855f

Compare with similar skills

Ax Audio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ax Audio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ax Audio this skilldosco/aithy107—~2.5kAutomated safety check: PassApache-2.0
Xsaimoeru-ai/airi50k1 repos~1.3kAutomated safety check: PassMIT
Whisper Speech RecognitionOrchestra-Research/AI-Research-SKILLs13k7 repos~1.9kAutomated safety check: NotesMIT
Openai Whisper APItrpc-group/trpc-agent-go1.9k12 repos~288Automated safety check: PassApache-2.0
Parakeet Sttsundial-org/awesome-openclaw-skills663—~771Automated safety check: PassNone
9Router Speech-to-Textdecolua/9router30k—~914Automated safety check: PassMIT

Similar skills

  • Xsai

    moeru-ai/airi

    A skill your agent uses when the user is building with xsai or any @xsai/ package, or is evaluating xsAI for a small OpenAI-compatible workflow with text generation, streaming, tool calling…

    50k GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    AI & LLM EngineeringAuto-check: notes
  • Openai Whisper API

    trpc-group/trpc-agent-go

    Transcribe audio via OpenAI Audio Transcriptions API (Whisper).

    1.9k GitHub starsUsed in 12 repos~288 tokens
    AI & LLM EngineeringAuto-check passed
  • Parakeet Stt

    sundial-org/awesome-openclaw-skills

    Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU).

    663 GitHub stars~771 tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    30k GitHub stars~914 tokensUpdated today
    Media & CreativeAuto-check passed
  • Openai Whisper API

    openclaw/openclaw

    OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

    392k GitHub starsUsed in 1 repo~518 tokens
    Media & CreativeAuto-check passed

More from dosco/aithy

All 17 skills in this repo
  • This skill helps an LLM generate correct AxAgent observability code using @ax-llm/ax.

    107 GitHub stars~4.4k tokensUpdated 1 mo ago
    Auto-check passed
  • This skill helps an LLM generate correct AxAgent tuning and evaluation code using @ax-llm/ax.

    107 GitHub stars~4.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Ax Gepa

    dosco/aithy

    This skill helps an LLM generate correct AxGEPA optimization code using @ax-llm/ax.

    107 GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Ax LLM

    dosco/aithy

    This skill helps with using the @ax-llm/ax TypeScript library for building LLM applications.

    107 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Ax MCP

    dosco/aithy

    This skill helps an LLM build correct native Model Context Protocol integrations with @ax-llm/ax.

    107 GitHub stars~4.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Ax Playbook

    dosco/aithy

    This skill helps an LLM generate correct playbook code using @ax-llm/ax.

    107 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Ax Audio

What does Ax Audio do?

This skill helps an LLM generate correct audio code with @ax-llm/ax. Ax Audio is an agent skill from dosco/aithy. This skill helps an LLM generate correct audio code with @ax-llm/ax.

When should I use Ax Audio?

Ax Audio fits situations like: the user asks about ai.transcribe(); signature audio inputs; agent audio behavior; .chat() conversational audio.

How do I install Ax Audio in Claude Code?

Run `npx skills add dosco/aithy --skill ax-audio -a claude-code`. Or copy the skill folder (.claude/skills/ax-audio in dosco/aithy) into .claude/skills/ax-audio in your project. Claude Code loads it when a task matches its description.

How do I install Ax Audio in Codex?

Run `npx skills add dosco/aithy --skill ax-audio -a codex`. Or copy the skill folder (.claude/skills/ax-audio in dosco/aithy) into .agents/skills/ax-audio in your project. Codex loads it when a task matches its description.

Can I use Ax Audio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dosco/aithy --skill ax-audio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ax-audio, .gemini/skills/ax-audio, .github/skills/ax-audio and .opencode/skills/ax-audio in your project.

What does Ax Audio need to run?

Going by SKILL.md and its folder, Ax Audio needs credentials named GROK_API_KEY. Our summary lists: A credential in GROK_API_KEY.

Does Ax Audio access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ax Audio safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ax Audio use?

Ax Audio is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ax Audio use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ax Audio?

Skills that share tags, products or a category with Ax Audio: Xsai (moeru-ai/airi, 50k stars), Whisper Speech Recognition (Orchestra-Research/AI-Research-SKILLs, 13k stars), Openai Whisper API (trpc-group/trpc-agent-go, 1.9k stars) and Parakeet Stt (sundial-org/awesome-openclaw-skills, 663 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ax Audio?

dosco (a GitHub user) maintains it in dosco/aithy, which has 107 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on August 31, 2026.

Source: dosco/aithy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.