AI Audio Models

k2-fsaOmniVoice
Free 6/1 hrafter ≈ $0.0055/use
Modalities
text / audio → audio
Released
Apr 1, 2026
Formats
WAV · MP3
Cloning
Zero-shot voice clone · Voice design

Massively multilingual zero-shot TTS with Voice Clone and Voice Design modes, prompt preprocessing, speed, duration, and diffusion controls.

ACE-StepACE-Step 1.5
Free 1/1 hrafter ≈ $0.132/use
Modalities
text → audio
Released
Jan 16, 2026
Duration
10–300 sec
Formats
MP3 · WAV · FLAC
Controls
BPM · Key · Lyrics · Style tags

Music model for structured songs with style tags, lyrics, BPM, key, language, and long durations.

Stability AIStable Audio 3 Medium
Free 2/1 hrafter ≈ $0.0366/use
Modalities
text / audio → audio
Released
Jun 10, 2024
Duration
Up to 180 sec (Stereo 44.1kHz)
Formats
WAV · MP3
Editing
Audio-to-audio · Inpainting

Fast high-quality music and sound generation with audio-to-audio editing, inpainting, and continuation.

Owen SongInflect Micro v2
Free 12/1 hrafter ≈ $0.0011/use
Modalities
text → audio
Released
Jul 25, 2026
Formats
WAV · MP3
Voice
Consistent English male narration

Instant CPU-efficient English narration with one consistent synthetic male voice.

Inworld AIInworld TTS-1.5 Mini
Free 12/1 hrafter ≈ $0.0275/1K chars
Modalities
text → audio
Released
Jan 21, 2026
Voices
20+ expressive character voices

Low-latency expressive text-to-speech optimized for real-time apps

xAIxAI Text-to-Speech
Free 12/1 hrafter ≈ $0.0165/1K chars
Modalities
text → audio
Released
Mar 16, 2026
Voices
20+ voices · Inline speech tags

Expressive text-to-speech with over two dozen voices, speech tags, and multilingual support

OpenAIOpenAI TTS (Text to Speech)
Free 8/3 hrafter ≈ $0.0165–0.0171/1K chars
Modalities
text → audio
Released
Nov 6, 2023
Formats
MP3 · AAC · Opus · FLAC
Voices
6 voices (Alloy, Shimmer, etc.)

Reliable OpenAI text-to-speech for clear narration, dialogue, and assistant voice output.

MiniMaxMiniMax Music 1.5
Upgrade required
Modalities
text → audio
Released
Aug 30, 2024
Formats
MP3 · WAV (up to 44.1kHz)
Structure
Lyrics · Verse/Chorus tags

Song model for natural vocals and rich arrangements with English or Chinese structured lyrics.

Fish AudioFish Audio S2.1 Pro
$15 / 1M
Modalities
text → audio
Released
Jun 1, 2026
Price
$15 / 1M
Formats
MP3 · WAV · FLAC · OGG

Flagship multilingual text-to-speech with natural language voice control and realtime streaming

Inworld AIInworld Realtime TTS-2
≈ $0.0385/1K chars
Modalities
text → audio
Released
May 5, 2026
Price
≈ $0.0385/1K chars
Formats
MP3 · WAV · FLAC · OGG

Conversational text-to-speech with realtime voice direction and audio-aware delivery

GoogleGemini 3.1 Flash TTS
$1.1 in · $22 out / 1M
Modalities
text → audio
Released
Apr 15, 2026
In / out price
$1.1 in · $22 out / 1M
Formats
MP3 · WAV · FLAC · OGG

Expressive text-to-speech with audio tags, multi-speaker dialogue, and 70+ languages

MiniMaxMiniMax Music 2.6
≈ $0.165/use
Modalities
text → audio
Released
Apr 10, 2026
Price
≈ $0.165/use
Formats
MP3 · WAV · FLAC · OGG

Promptable full-song generation with vocals, lyrics, BPM and key control

MiniMaxMiniMax Music Cover
≈ $0.165/use
Modalities
text / audio → audio
Released
Apr 10, 2026
Price
≈ $0.165/use
Formats
MP3 · WAV · FLAC · OGG

Audio-to-audio song transformation that preserves melody while changing style

RunwareACE-Step v1.5 XL Base
≈ $0.00028–0.0003/sec
Modalities
text / audio → audio
Released
Apr 2, 2026
Price
≈ $0.00028–0.0003/sec
Duration
30–300 sec
Formats
MP3 · WAV · FLAC · OGG

4B music generation model with higher audio quality and full editing task support

RunwareACE-Step v1.5 XL SFT
≈ $0.000143–0.000176/sec
Modalities
text / audio → audio
Released
Apr 2, 2026
Price
≈ $0.000143–0.000176/sec
Duration
30–300 sec
Formats
MP3 · WAV · FLAC · OGG

Highest-quality 4B music generation model with CFG-controlled prompt adherence

RunwareACE-Step v1.5 XL Turbo
≈ $0.0000165–0.000033/sec
Modalities
text / audio → audio
Released
Apr 2, 2026
Price
≈ $0.0000165–0.000033/sec
Duration
30–300 sec
Formats
MP3 · WAV · FLAC · OGG

Fast 4B music generation model with 8-step inference for higher-quality rapid iteration

MiniMaxMiniMax Speech 2.8
≈ $0.066–0.11/1K chars
Modalities
text → audio
Released
Jan 29, 2026
Price
≈ $0.066–0.11/1K chars
Formats
MP3 · WAV · FLAC · OGG

High-quality text-to-speech with expressive, natural voice synthesis

Inworld AIInworld TTS-1.5 Max
≈ $0.055/1K chars
Modalities
text → audio
Released
Jan 21, 2026
Price
≈ $0.055/1K chars
Formats
MP3 · WAV · FLAC · OGG

High-fidelity expressive text-to-speech with rich prosody and multilingual support

AlibabaQwen3-TTS 1.7B Base
≈ $0.0165/1K chars
Modalities
text / audio → audio
Released
Jan 1, 2026
Price
≈ $0.0165/1K chars
Formats
MP3 · WAV · FLAC · OGG

High-quality multilingual text-to-speech with voice cloning and ultra-low latency

AlibabaQwen3-TTS 1.7B CustomVoice
≈ $0.0165/1K chars
Modalities
text → audio
Released
Jan 1, 2026
Price
≈ $0.0165/1K chars
Formats
MP3 · WAV · FLAC · OGG

Text-to-speech with preset premium timbres and precise style control

AlibabaQwen3-TTS 1.7B VoiceDesign
≈ $0.0165/1K chars
Modalities
text → audio
Released
Jan 1, 2026
Price
≈ $0.0165/1K chars
Formats
MP3 · WAV · FLAC · OGG

Text-to-speech with voice creation from natural language descriptions

RunwareACE-Step v1.5 Base
≈ $0.000165–0.0099/use
Modalities
text / audio → audio
Released
Jan 1, 2026
Price
≈ $0.000165–0.0099/use
Duration
30–300 sec
Formats
MP3 · WAV · FLAC · OGG

Open-source music generation with voice cloning, lyric editing, and multilingual support

RunwareACE-Step v1.5 Turbo
≈ $0.00011–0.0066/use
Modalities
text / audio → audio
Released
Jan 1, 2026
Price
≈ $0.00011–0.0066/use
Duration
30–300 sec
Formats
MP3 · WAV · FLAC · OGG

Fast music generation optimized for speed with reduced inference steps

RunwareDia2 2B
Usage based
Modalities
text → audio
Released
Nov 19, 2025
Price
Usage based
Formats
MP3 · WAV · FLAC · OGG

Streaming dialogue TTS with voice cloning, non-verbal cues, and multi-speaker support