Xsai
moeru-ai/airi
A skill your agent uses when the user is building with xsai or any @xsai/ package, or is evaluating xsAI for a small OpenAI-compatible workflow with text generation, streaming, tool calling…
This skill helps an LLM generate correct audio code with @ax-llm/ax.
$ npx skills add dosco/aithy --skill ax-audio -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install dosco/aithy ax-audio --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ax-audio .claude/skills/ax-audio && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ax-audio" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-audio into .claude/skills/ax-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-audio", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/dosco/aithy/tree/main/.claude/skills/ax-audioType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add dosco/aithy --skill ax-audio -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install dosco/aithy ax-audio --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/ax-audio .agents/skills/ax-audio && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ax-audio" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-audio into .agents/skills/ax-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-audio", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dosco/aithy --skill ax-audio -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install dosco/aithy ax-audio --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/ax-audio .cursor/skills/ax-audio && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ax-audio" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-audio into .cursor/skills/ax-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-audio", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/dosco/aithy.git --path .claude/skills/ax-audio--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add dosco/aithy --skill ax-audio -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install dosco/aithy ax-audio --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/ax-audio .gemini/skills/ax-audio && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ax-audio" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-audio into .gemini/skills/ax-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-audio", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install dosco/aithy ax-audioInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add dosco/aithy --skill ax-audio -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/ax-audio .github/skills/ax-audio && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ax-audio" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-audio into .github/skills/ax-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-audio", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dosco/aithy --skill ax-audio -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install dosco/aithy ax-audio --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dosco/aithy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/ax-audio .opencode/skills/ax-audio && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ax-audio" agent skill from https://github.com/dosco/aithy/tree/main/.claude/skills/ax-audio into .opencode/skills/ax-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ax-audio", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ax-audioThis skill helps an LLM generate correct audio code with @ax-llm/ax.
Ax Audio is an agent skill from dosco/aithy. This skill helps an LLM generate correct audio code with @ax-llm/ax. Use when the user asks about ai.transcribe(), ai.speak(), signature audio inputs or outputs, agent audio behavior, .chat() conversational audio, OpenAI audio or realtime models, Gemini Live native audio, Grok Voice Agent models, voices, formats, transcripts, or how audio fits with structured outputs.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Transcription, Structured output and tool calling and Speech recognition and synthesis. It works with OpenAI. The repository describes itself as: A personal AI agent that can work safely on your machine, remember useful context, and keep its data under your control. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 0c9855f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GROK_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ax Audio loads about 2.5k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 547 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from dosco/aithy at commit 0c9855f, republished under its Apache-2.0 licence (© dosco). 547 words, ~2,506 tokens.
.claude/skills/ax-audio/SKILL.md (or your agent's skills folder).Use this skill for audio in Ax. Pick the smallest audio surface that matches the job:
ai.transcribe(...) for batch speech-to-text.ai.speak(...) for batch text-to-speech.speech:audio signature outputs for structured programs that should return synthesized audio artifacts..chat() audio config for conversational or realtime audio turns.:audio is an audio input value: { data, format?, mimeType?, sampleRate?, channels? }.:audio is a scripted audio artifact. The model returns plain text for that field; Ax synthesizes it after structured output parsing.string, not a binary object..chat() and modelConfig.audio.speech options, not modelConfig.audio.import { ai } from '@ax-llm/ax';
const llm = ai({ name: 'openai', apiKey: process.env.OPENAI_APIKEY! });
const transcript = await llm.transcribe({
audio: { data: base64Wav, format: 'wav' },
model: 'gpt-4o-mini-transcribe',
language: 'en',
prompt: 'Product support call',
});
const speech = await llm.speak({
text: transcript.text,
model: 'gpt-4o-mini-tts',
voice: 'alloy',
format: 'mp3',
});
console.log(transcript.text);
console.log(speech.data);
console.log(speech.transcript);Providers without the requested batch audio capability throw AxMediaNotSupportedError.
import { ai, ax } from '@ax-llm/ax';
const llm = ai({ name: 'openai', apiKey: process.env.OPENAI_APIKEY! });
const say = ax('question:string -> speech:audio, summary:string');
const result = await say.forward(
llm,
{ question: 'Explain retries in one sentence.' },
{
speech: {
speak: { voice: 'alloy', format: 'mp3' },
fields: {
speech: { voice: 'alloy' },
},
},
}
);
console.log(result.summary);
console.log(result.speech.data);
console.log(result.speech.mimeType);
console.log(result.speech.transcript);The model emits a text script for speech; Ax replaces it with AxChatAudioOutput after result selection. If the field already contains an audio artifact with { data } or { id }, Ax leaves it alone.
import { agent, ai } from '@ax-llm/ax';
const llm = ai({ name: 'openai', apiKey: process.env.OPENAI_APIKEY! });
const voiceAgent = agent(
'recording:audio, question:string -> speech:audio, summary:string',
{
agentIdentity: {
name: 'Voice Assistant',
description: 'Answers spoken requests with spoken and written output',
},
contextFields: [],
}
);
const result = await voiceAgent.forward(
llm,
{
recording: { data: base64Wav, format: 'wav' },
question: 'What should I do next?',
},
{
speech: {
transcribe: { model: 'gpt-4o-mini-transcribe' },
speak: { voice: 'alloy', format: 'mp3' },
},
}
);
console.log(result.summary);
console.log(result.speech.data);The agent runtime transcribes recording first and passes the transcript through the internal agent stages. Use direct ax(...) or .chat() when you specifically want native audio understanding in the model call.
.chat() AudioUse modelConfig.audio for conversational audio turns where audio is part of the chat response instead of a structured signature field.
const res = await llm.chat({
chatPrompt: [{ role: 'user', content: 'Say hello out loud.' }],
modelConfig: {
audio: { output: { enabled: true, voice: 'alloy', format: 'wav' } },
},
});
console.log(res.results[0]?.content);
console.log(res.results[0]?.audio?.data);
console.log(res.results[0]?.audio?.transcript);type AxAudioFormat =
| 'wav'
| 'mp3'
| 'flac'
| 'opus'
| 'aac'
| 'pcm16'
| 'pcm'
| 'ogg'
| 'raw'
| 'mulaw'
| 'ulaw'
| 'alaw';
type AxSpeechConfig = {
transcribe?: {
model?: string;
language?: string;
prompt?: string;
};
speak?: {
model?: string;
voice?: string;
format?: AxAudioFormat;
};
fields?: Record<
string,
{
model?: string;
voice?: string;
format?: AxAudioFormat;
}
>;
};Use axAIOpenAIAudioDefaultConfig() for OpenAI request-based audio chat:
gpt-audio-minialloywavwav, mp3wav, mp3, flac, opus, aac, pcm16import { ai, axAIOpenAIAudioDefaultConfig } from '@ax-llm/ax';
const openai = ai({
name: 'openai',
apiKey: process.env.OPENAI_APIKEY!,
config: axAIOpenAIAudioDefaultConfig(),
});
const res = await openai.chat({
chatPrompt: [
{
role: 'user',
content: [
{ type: 'text', text: 'What is in this recording?' },
{ type: 'audio', data: base64Wav, format: 'wav' },
],
},
],
});
console.log(res.results[0]?.content);
console.log(res.results[0]?.audio?.data);Use axAIOpenAIRealtimeDefaultConfig() for OpenAI realtime speech-to-speech:
gpt-realtime-2marinpcm16audio/pcm, mono, 24000 Hz30000Use axAIOpenAIRealtimeTranscriptionDefaultConfig() for realtime transcript deltas:
gpt-realtime-whisperaudio/pcm, mono, 24000 HzcontentRealtime models use a one-turn WebSocket call under .chat(). In Node, pass a WebSocket constructor through request options:
import WebSocket from 'ws';
import { ai, axAIOpenAIRealtimeDefaultConfig } from '@ax-llm/ax';
const openai = ai({
name: 'openai',
apiKey: process.env.OPENAI_APIKEY!,
config: axAIOpenAIRealtimeDefaultConfig(),
});
const stream = await openai.chat(
{
chatPrompt: [{ role: 'user', content: 'Say hello out loud.' }],
},
{ stream: true, webSocket: WebSocket }
);For follow-up turns, keep the assistant audio reference in history:
await openai.chat({
chatPrompt: [
{ role: 'assistant', audio: { id: previousAudioId } },
{ role: 'user', content: 'Repeat that more slowly.' },
],
});Use axAIGoogleGeminiLiveAudioDefaultConfig() for Gemini native audio:
gemini-2.5-flash-native-audio-preview-12-2025Korepcm1624000audio/pcm;rate=16000, mono30000import { ai, axAIGoogleGeminiLiveAudioDefaultConfig } from '@ax-llm/ax';
const gemini = ai({
name: 'google-gemini',
apiKey: process.env.GOOGLE_APIKEY!,
config: axAIGoogleGeminiLiveAudioDefaultConfig(),
});
const res = await gemini.chat({
chatPrompt: [
{
role: 'user',
content: [
{ type: 'text', text: 'Answer this spoken question.' },
{
type: 'audio',
data: base64Pcm16,
format: 'pcm16',
sampleRate: 16000,
channels: 1,
},
],
},
],
});
console.log(res.results[0]?.content);
console.log(res.results[0]?.audio?.data);Gemini Live uses a one-turn WebSocket call under .chat(). It expects PCM input for native audio turns; use format: 'pcm16' or mimeType: 'audio/pcm;rate=16000'.
Use axAIGrokVoiceDefaultConfig() for xAI Grok Voice Agent:
grok-voice-think-fast-1.0evepcm1624000audio/pcm, mono, 24000 Hz30000import WebSocket from 'ws';
import { ai, axAIGrokVoiceDefaultConfig } from '@ax-llm/ax';
const grok = ai({
name: 'grok',
apiKey: process.env.GROK_API_KEY!,
config: axAIGrokVoiceDefaultConfig(),
});
const res = await grok.chat(
{
chatPrompt: [{ role: 'user', content: 'Say hello out loud.' }],
},
{ webSocket: WebSocket }
);
console.log(res.results[0]?.content);
console.log(res.results[0]?.audio?.data);Grok Voice uses a one-turn WebSocket call under .chat(). It expects PCM input for spoken input turns; use format: 'pcm16' or mimeType: 'audio/pcm'.
OpenAI audio chat, OpenAI Realtime, Gemini Live, and Grok Voice all default to non-streaming, but each can stream deltas when you pass { stream: true }.
const stream = await llm.chat(
{
chatPrompt: [{ role: 'user', content: 'Say hello.' }],
},
{ stream: true }
);
for await (const chunk of stream) {
const audio = chunk.results[0]?.audio;
if (audio?.isDelta) {
playAudioChunk(audio.data);
}
}Use signature audio outputs for structured speech artifacts:
const gen = ax('question:string -> answer:string, speech:audio');Use .chat() audio when the response itself is a conversational audio turn. Do not combine .chat() audio output with provider-native structured response formats unless that provider explicitly supports the combination.
© dosco, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/ax-audio of dosco/aithy.
Open the folder on GitHubat commit 0c9855f
Ax Audio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ax Audio this skilldosco/aithy | 107 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Xsaimoeru-ai/airi | 50k | 1 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Whisper Speech RecognitionOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~1.9k | Automated safety check: Notes | MIT | |
| Openai Whisper APItrpc-group/trpc-agent-go | 1.9k | 12 repos | ~288 | Automated safety check: Pass | Apache-2.0 | |
| Parakeet Sttsundial-org/awesome-openclaw-skills | 663 | — | ~771 | Automated safety check: Pass | None | |
| 9Router Speech-to-Textdecolua/9router | 30k | — | ~914 | Automated safety check: Pass | MIT |
moeru-ai/airi
A skill your agent uses when the user is building with xsai or any @xsai/ package, or is evaluating xsAI for a small OpenAI-compatible workflow with text generation, streaming, tool calling…
Orchestra-Research/AI-Research-SKILLs
Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.
trpc-group/trpc-agent-go
Transcribe audio via OpenAI Audio Transcriptions API (Whisper).
sundial-org/awesome-openclaw-skills
Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU).
decolua/9router
Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.
openclaw/openclaw
OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.
dosco/aithy
This skill helps an LLM generate correct AxAgent observability code using @ax-llm/ax.
dosco/aithy
This skill helps an LLM generate correct AxAgent tuning and evaluation code using @ax-llm/ax.
dosco/aithy
This skill helps an LLM generate correct AxGEPA optimization code using @ax-llm/ax.
dosco/aithy
This skill helps with using the @ax-llm/ax TypeScript library for building LLM applications.
dosco/aithy
This skill helps an LLM build correct native Model Context Protocol integrations with @ax-llm/ax.
dosco/aithy
This skill helps an LLM generate correct playbook code using @ax-llm/ax.
Works with
Categories
This skill helps an LLM generate correct audio code with @ax-llm/ax. Ax Audio is an agent skill from dosco/aithy. This skill helps an LLM generate correct audio code with @ax-llm/ax.
Ax Audio fits situations like: the user asks about ai.transcribe(); signature audio inputs; agent audio behavior; .chat() conversational audio.
Run `npx skills add dosco/aithy --skill ax-audio -a claude-code`. Or copy the skill folder (.claude/skills/ax-audio in dosco/aithy) into .claude/skills/ax-audio in your project. Claude Code loads it when a task matches its description.
Run `npx skills add dosco/aithy --skill ax-audio -a codex`. Or copy the skill folder (.claude/skills/ax-audio in dosco/aithy) into .agents/skills/ax-audio in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dosco/aithy --skill ax-audio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ax-audio, .gemini/skills/ax-audio, .github/skills/ax-audio and .opencode/skills/ax-audio in your project.
Going by SKILL.md and its folder, Ax Audio needs credentials named GROK_API_KEY. Our summary lists: A credential in GROK_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ax Audio is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ax Audio: Xsai (moeru-ai/airi, 50k stars), Whisper Speech Recognition (Orchestra-Research/AI-Research-SKILLs, 13k stars), Openai Whisper API (trpc-group/trpc-agent-go, 1.9k stars) and Parakeet Stt (sundial-org/awesome-openclaw-skills, 663 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
dosco (a GitHub user) maintains it in dosco/aithy, which has 107 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on August 31, 2026.
Source: dosco/aithy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.