AI Speech-to-Text Models

NVIDIANemotron 3.5 ASR Streaming 0.6B
Free 6/1 hrafter ≈ $0.033/use
Modalities
audio → text
Released
Dec 24, 2024
Input
Audio / speech file
Output
Streaming transcription + timestamps

⚡️ Low-latency multilingual transcription | streaming ASR architecture | punctuation | capitalization | auto language detection

AlibabaSenseVoiceSmall
Free 6/1 hrafter ≈ $0.011/use
Modalities
audio → text
Released
Jul 3, 2024
Input
Audio / speech file
Output
Multilingual transcription + emotions

Fast multilingual transcription with FunASR SenseVoice-Small.

OpenAIWhisper v3 Diarization
Free 1/1 hrafter ≈ $0.132/use
Modalities
audio → text
Released
Nov 14, 2023
Input
Audio / speech file
Output
Speaker diarization + timestamps

⚡️ Fast audio transcription | whisper v3 | speaker diarization | word level timestamps | prompt

NVIDIANemotron 3.5 ASR Streaming Multilingual 0.6B
≈ $0.00000386/sec
Modalities
audio → text
Released
Aug 13, 2026
Price
≈ $0.00000386/sec

Nemotron 3.5 ASR Streaming Multilingual 0.6B is a speech recognition model from NVIDIA. Its prompt-conditioned, cache-aware FastConformer-RNNT design targets low-latency transcription across more than 40 languages for real-time captioning, voice...

Mistral AIVoxtral Small 24B 2507 STT
≈ $0.000058/sec
Modalities
audio → text
Released
Aug 13, 2026
Price
≈ $0.000058/sec

Voxtral Small 24B 2507 STT is a speech transcription model from Mistral AI. It is suited for transcription, translation, and audio understanding workloads that benefit from its larger model capacity.

Mistral AIVoxtral Mini 3B 2507
≈ $0.0000193/sec
Modalities
audio → text
Released
Aug 13, 2026
Price
≈ $0.0000193/sec

Voxtral Mini 3B 2507 is a speech and audio understanding model from Mistral AI. It is suited for transcription, translation, and compact audio processing workloads.

AlibabaQwen3 ASR 1.7B
≈ $0.0000087/sec
Modalities
audio → text
Released
Aug 13, 2026
Price
≈ $0.0000087/sec

Qwen3 ASR 1.7B is an automatic speech recognition model from Qwen. It supports multilingual language identification and transcription across 30 languages and 22 Chinese dialects, with streaming and offline inference...

AlibabaQwen3 ASR 0.6B
≈ $0.00000386/sec
Modalities
audio → text
Released
Aug 13, 2026
Price
≈ $0.00000386/sec

Qwen3 ASR 0.6B is a compact automatic speech recognition model from Qwen. It supports multilingual language identification and transcription across 30 languages and 22 Chinese dialects, with streaming and offline...

OpenAIGPT Transcribe
≈ $0.000087/sec
Modalities
audio → text
Released
Aug 5, 2026
Price
≈ $0.000087/sec

GPT Transcribe is a high-accuracy speech-to-text model from OpenAI. It is suited for recorded audio, streamed file transcription, and committed Realtime turns, with free-form context, keyword hints, and multiple language...

Fish AudioTranscribe 1
≈ $0.000116/sec
Modalities
audio → text
Released
Jul 29, 2026
Price
≈ $0.000116/sec

Transcribe 1 is a speech-to-text model from Fish Audio. It is suited for audio transcription with automatic language detection and can return timestamped word-level segments when alignment details are requested.

xAIGrok STT 1.0
≈ $0.0000322/sec
Modalities
audio → text
Released
Jul 23, 2026
Price
≈ $0.0000322/sec

Grok STT is SpaceXAI's speech-to-text model, available via the REST /v1/stt endpoint. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio.

DeepgramNova-3
≈ $0.0000832–0.000101/sec
Modalities
audio → text
Released
Jul 15, 2026
Price
≈ $0.0000832–0.000101/sec

Deepgram Nova-3 general-purpose speech-to-text model with monolingual and multilingual transcription support.

MicrosoftMAI-Transcribe 1.5
≈ $0.000116/sec
Modalities
audio → text
Released
Jun 2, 2026
Price
≈ $0.000116/sec

MAI-Transcribe 1.5 is a multilingual speech-to-text model from Microsoft AI. It is suited for captions, call transcription, subtitling, accessibility, and other voice-enabled applications, with reliable transcription across 43 languages, diverse...

NVIDIAParakeet TDT 0.6B v3
≈ $0.000029/sec
Modalities
audio → text
Released
May 27, 2026
Price
≈ $0.000029/sec

Parakeet TDT 0.6B v3 is NVIDIA's 600M-parameter multilingual speech-to-text model built on the FastConformer-TDT architecture. Trained on the Granary dataset (670,000+ hours of audio), it supports automatic language detection across...

Mistral AIVoxtral Mini Transcribe
≈ $0.000058/sec
Modalities
audio → text
Released
May 15, 2026
Price
≈ $0.000058/sec

Voxtral Mini Transcribe is Mistral's speech-to-text model, derived from the Voxtral Mini family. It accepts audio input and returns transcribed text via the standard transcription API. Suited for transcribing meetings,...

AlibabaQwen3 ASR Flash
≈ $0.0000406/sec
Modalities
audio → text
Released
May 14, 2026
Price
≈ $0.0000406/sec

Qwen3-ASR-Flash is Alibaba's automatic speech recognition service, built on the Qwen3-Omni foundation and trained on tens of millions of hours of multimodal speech data. The model handles 11 languages —...

OpenAIgpt-realtime-whisper
≈ $0.000324/sec
Modalities
audio → text
Released
May 7, 2026
Price
≈ $0.000324/sec

GPT Realtime Whisper is a streaming speech-to-text model for applications that need low-latency transcript deltas from live audio. It is designed for realtime use cases where developers need to tune latency and accuracy. GPT Realtime Whisper is priced by audio duration rather than text tokens.

GoogleChirp 3
≈ $0.000309/sec
Modalities
audio → text
Released
May 5, 2026
Price
≈ $0.000309/sec

Chirp 3 is Google's latest multilingual speech-to-text model. It offers enhanced transcription accuracy across 24 GA languages and 77+ preview languages, with support for automatic language detection, automatic punctuation, and...

OpenAIGPT-4o Mini Transcribe
$1.42 in · $5.7 out / 1M
Modalities
audio → text
Released
May 1, 2026
In / out price
$1.42 in · $5.7 out / 1M

GPT-4o Mini Transcribe is OpenAI's smaller, cost-efficient speech-to-text model built on GPT-4o Mini audio capabilities. It's priced per token (input and output), making it suitable for high-volume transcription workflows that...

OpenAIWhisper Large V3
≈ $0.0000087/sec
Modalities
audio → text
Released
May 1, 2026
Price
≈ $0.0000087/sec

Whisper Large V3 is OpenAI's open-source automatic speech recognition model offering both audio transcription and translation. It supports 99+ languages and accepts common audio formats including mp3, mp4, wav, webm,...

OpenAIWhisper Large V3 Turbo
≈ $0.00000386/sec
Modalities
audio → text
Released
May 1, 2026
Price
≈ $0.00000386/sec

Whisper Large V3 Turbo is an optimized version of OpenAI's Whisper Large V3 speech recognition model, designed for speed and cost efficiency. It supports transcription across 99+ languages with a...

OpenAIWhisper 1
≈ $0.000114–0.000116/sec
Modalities
audio → text
Released
Apr 27, 2026
Price
≈ $0.000114–0.000116/sec

Whisper is OpenAI's open-source automatic speech recognition model, available via API as `whisper-1`. It supports transcription and translation across 50+ languages from audio files up to 25 MB. Accepts formats...

OpenAIGPT-4o Transcribe
$2.85 in · $11.4 out / 1M
Modalities
audio → text
Released
Apr 27, 2026
In / out price
$2.85 in · $11.4 out / 1M

GPT-4o Transcribe is OpenAI's high-quality speech-to-text model built on GPT-4o audio capabilities. It's priced per token (input and output), making it suitable for workflows that benefit from token-level billing transparency.

xAIGrok STT
≈ $0.0000319/sec
Modalities
audio → text
Released
Mar 16, 2026
Price
≈ $0.0000319/sec

Transcribe audio to text in 25 languages with batch and streaming modes.