AI Audio Models

ACE-StepACE-Step 1.5
1/1 hr included

Music model for structured songs with style tags, lyrics, BPM, key, language, and long durations.

k2-fsaOmniVoice
6/1 hr included

Massively multilingual zero-shot TTS with Voice Clone and Voice Design modes, prompt preprocessing, speed, duration, and diffusion controls.

Stability AIStable Audio 3 Medium
2/1 hr included

Fast high-quality music and sound generation with audio-to-audio editing, inpainting, and continuation.

MMAudioMMAudio V2 (Video to Audio)
1/4 hr included

Video-to-audio model that adds synchronized ambience, Foley, and effects from video plus prompt.

Sony ResearchWoosh (Text/Video to Audio)
1/4 hr included

Sony video/audio model for fast text sound effects or synchronized audio from video.

Owen SongInflect Micro v2
12/1 hr included

Instant CPU-efficient English narration with one consistent synthetic male voice.

HexgradKokoro 82M (English TTS)
12/1 hr included

Small English TTS model with many voices and natural American, British, and international accents.

AlibabaQwen3 TTS Custom Voice
2/1 hr includedafter ≈ $0.0165/1K chars
Modalities
text → audio
Released
Jan 1, 2026
Formats
MP3 · WAV · FLAC · OGG

Text-to-speech with preset premium timbres and precise style control

AlibabaQwen3 TTS Voice Design
1/1 hr includedafter ≈ $0.0165/1K chars
Modalities
text → audio
Released
Jan 1, 2026
Formats
MP3 · WAV · FLAC · OGG

Text-to-speech with voice creation from natural language descriptions

Inworld AIInworld TTS-1.5 Mini
12/1 hr includedafter ≈ $0.0275/1K chars
Modalities
text → audio
Released
Jan 21, 2026
Formats
MP3 · WAV · FLAC · OGG

Low-latency expressive text-to-speech optimized for real-time apps

xAIxAI Text-to-Speech
12/1 hr includedafter ≈ $0.0165/1K chars
Modalities
text → audio
Released
Mar 16, 2026
Formats
MP3 · WAV · FLAC · OGG

Expressive text-to-speech with five voices, speech tags, and multilingual support

OpenAIOpenAI TTS (Text to Speech)
8/3 hr includedafter ≈ $0.0165/1K chars
Modalities
text → audio
Released
Nov 6, 2023

Reliable OpenAI text-to-speech for clear narration, dialogue, and assistant voice output.

MiniMaxMiniMax Music 1.5
Upgrade required

Song model for natural vocals and rich arrangements with English or Chinese structured lyrics.

fish-audioS1
≈ $0.0174/1K chars
Released
Jul 29, 2026
Price
≈ $0.0174/1K chars

S1 is a multilingual text-to-speech model from Fish Audio. It is suited for voice applications that need broad emotional expression, using parenthetical controls to guide speaking style across its supported...

fish-audioS2 Pro
≈ $0.0174/1K chars
Released
Jul 29, 2026
Price
≈ $0.0174/1K chars

S2 Pro is a multilingual text-to-speech model from Fish Audio. It is suited for expressive narration and multi-speaker dialogue, with natural-language controls for speaking style and emotion.

MicrosoftMAI-Voice-2-Flash
≈ $0.0174/1K chars
Released
Jul 23, 2026
Price
≈ $0.0174/1K chars

MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15 languages...

AlibabaQwen-Audio-3.0-TTS Flash
≈ $0.0174/1K chars
Released
Jul 23, 2026
Price
≈ $0.0174/1K chars

Qwen-Audio-3.0-TTS Flash is Alibaba's fast, cost-efficient text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

AlibabaQwen-Audio-3.0-TTS Plus
≈ $0.0232/1K chars
Released
Jul 23, 2026
Price
≈ $0.0232/1K chars

Qwen-Audio-3.0-TTS Plus is Alibaba's higher-quality text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

deepgramAura-2
≈ $0.0348/1K chars
Released
Jul 16, 2026
Price
≈ $0.0348/1K chars

Aura-2 is a multilingual text-to-speech model from Deepgram. It supports Deepgram’s canonical Aura-2 voice catalog for speech synthesis across multiple languages.

MiniMaxSpeech 2.8 HD
≈ $0.116/1K chars
Released
Jul 16, 2026
Price
≈ $0.116/1K chars

MiniMax Speech 2.8 HD is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs.

MiniMaxSpeech 2.8 Turbo
≈ $0.0696/1K chars
Released
Jul 16, 2026
Price
≈ $0.0696/1K chars

MiniMax Speech 2.8 Turbo is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs.

MicrosoftMAI-Voice-2
≈ $0.0255/1K chars
Released
Jun 2, 2026
Price
≈ $0.0255/1K chars

MAI-Voice-2 is an expressive text-to-speech model from Microsoft. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18 locales,...

Fish AudioFish Audio S2.1 Pro
≈ $0.0174/1K chars
Modalities
text → audio
Released
Jun 1, 2026
Price
≈ $0.0174/1K chars
Formats
MP3 · WAV · FLAC · OGG

Flagship multilingual text-to-speech with natural language voice control and realtime streaming

Inworld AIInworld Realtime TTS-2
≈ $0.0385/1K chars
Modalities
text → audio
Released
May 5, 2026
Price
≈ $0.0385/1K chars
Formats
MP3 · WAV · FLAC · OGG

Conversational text-to-speech with realtime voice direction and audio-aware delivery