Audio models
Speech model prices
Transcription, speech synthesis and realtime voice models, with the price each host lists. No quality ranking: every benchmark on this site measures text. Language model prices.
| Whisper | OpenAI | transcription | — | free | $0.003 | free | 2023-09-01 |
| GPT-4o Mini Transcribe (OpenAI) | OpenAI | transcription | — | $1.25 | $1.25 | $5.00 | 2025-03-20 |
| Gemini 3.5 Transcribe | transcription | — | $2 | $2.00 | $12.00 | 2026-08-26 | |
| GPT-4o Transcribe (OpenAI) | OpenAI | transcription | — | $2.5 | $2.50 | $10.00 | 2025-03-20 |
| gpt-4o-transcribe-diarize | OpenAI | transcription | — | $2.5 | $2.50 | $10.00 | — |
| Gemini 3.1 Flash Live Preview | realtime | — | $3 | $0.75 | $4.50 | 2026-03-26 | |
| gemini-2-5-flash-native-audio | realtime | — | $3 | $0.50 | $2.00 | — | |
| gemini-2-5-flash-native-audio-preview-09 | realtime | — | $3 | $0.50 | $2.00 | — | |
| gemini-2-5-flash-native-audio-preview-12 | realtime | — | $3 | $0.50 | $2.00 | — | |
| gemini-live-2-5-flash-native-audio | realtime | — | $3 | $0.50 | $2.00 | — | |
| gemini-live-2-5-flash-preview-native-audio-09 | realtime | — | $3 | $0.50 | $2.00 | — | |
| nova-2-sonic | Amazon | realtime | — | $3 | $0.33 | $2.75 | — |
| nova-sonic | Amazon | realtime | — | $3.4 | $0.06 | $0.24 | — |
| Gemini 3.5 Live Translate Preview | realtime | — | $3.5 | $3.50 | $21.00 | 2026-06-09 | |
| Gemini 3.5 Transcribe Live | transcription | — | $3.5 | $3.50 | $21.00 | 2026-08-26 | |
| GPT-Realtime mini | OpenAI | realtime | — | $10 | $0.60 | $2.40 | 2025-10-10 |
| gpt-realtime-2-1-mini | OpenAI | realtime | — | $10 | $0.60 | $2.40 | — |
| gpt-4o-mini-realtime | OpenAI | realtime | — | $11 | $0.66 | $2.64 | — |
| gpt-realtime | OpenAI | realtime | — | $32 | $4.00 | $16.00 | — |
| GPT-Realtime-1.5 | OpenAI | realtime | — | $32 | $4.00 | $16.00 | 2026-02-23 |
| gpt-realtime-2 | OpenAI | realtime | — | $32 | $4.00 | $24.00 | 2026-05-07 |
| gpt-realtime-2.1 | OpenAI | realtime | — | $32 | $4.00 | $24.00 | 2026-07-06 |
| gpt-4o-realtime | OpenAI | realtime | — | $110 | $5.50 | $22.00 | — |
| asr | — | transcription | — | — | free | free | — |
| asr-large | — | transcription | — | — | free | free | — |
| base | — | transcription | — | — | — | — | — |
| base-conversationalai | — | transcription | — | — | — | — | — |
| base-finance | — | transcription | — | — | — | — | — |
| base-general | — | transcription | — | — | — | — | — |
| base-meeting | — | transcription | — | — | — | — | — |
| base-phonecall | — | transcription | — | — | — | — | — |
| base-video | — | transcription | — | — | — | — | — |
| base-voicemail | — | transcription | — | — | — | — | — |
| best | — | transcription | — | — | — | — | — |
| Canopy Labs Orpheus Arabic Saudi | — | speech | — | — | — | — | 2025-12-16 |
| Canopy Labs Orpheus V1 English | — | speech | — | — | — | — | 2025-12-19 |
| chirp | speech | — | — | — | — | — | |
| chirp-3 | transcription | — | — | — | — | — | |
| eleven-multilingual-v2 | — | speech | — | — | — | — | — |
| eleven-v3 | — | speech | — | — | — | — | — |
| enhanced | — | transcription | — | — | — | — | — |
| enhanced-finance | — | transcription | — | — | — | — | — |
| enhanced-general | — | transcription | — | — | — | — | — |
| enhanced-meeting | — | transcription | — | — | — | — | — |
| enhanced-phonecall | — | transcription | — | — | — | — | — |
| Gemini 2.5 Flash Preview TTS | speech | — | — | $0.50 | $10.00 | 2025-05-01 | |
| Gemini 3.1 Flash TTS Preview | speech | — | — | free | free | 2026-04-15 | |
| generative | — | speech | — | — | — | — | — |
| gpt-4o-mini-tts | OpenAI | speech | — | — | $2.50 | $10.00 | — |
| gpt-live-1 | OpenAI | realtime | — | — | — | — | — |
| gpt-live-transcribe | OpenAI | transcription | — | — | — | — | — |
| gpt-realtime-translate | OpenAI | realtime | — | — | — | — | — |
| gpt-realtime-whisper | OpenAI | transcription | — | — | — | — | 2026-05-07 |
| gpt-transcribe | OpenAI | transcription | — | — | — | — | — |
| long-form | — | speech | — | — | — | — | — |
| Lyria 3 Clip Preview | chat | $0.040 | — | free | free | 2026-03-25 | |
| Lyria 3 Pro Preview | chat | $0.080 | — | free | free | 2026-03-25 | |
| lyria-002 | speech | $0.060 | — | — | — | — | |
| muse-voice-transcribe-1-0 | Meta | transcription | — | — | — | — | — |
| nano | — | transcription | — | — | — | — | — |
| neural | — | speech | — | — | — | — | — |
| nova | Amazon | transcription | — | — | — | — | — |
| nova-2 | Amazon | transcription | — | — | — | — | — |
| nova-2-atc | Amazon | transcription | — | — | — | — | — |
| nova-2-automotive | Amazon | transcription | — | — | — | — | — |
| nova-2-conversationalai | Amazon | transcription | — | — | — | — | — |
| nova-2-drivethru | Amazon | transcription | — | — | — | — | — |
| nova-2-finance | Amazon | transcription | — | — | — | — | — |
| nova-2-general | Amazon | transcription | — | — | — | — | — |
| nova-2-meeting | Amazon | transcription | — | — | — | — | — |
| nova-2-phonecall | Amazon | transcription | — | — | — | — | — |
| nova-2-video | Amazon | transcription | — | — | — | — | — |
| nova-2-voicemail | Amazon | transcription | — | — | — | — | — |
| nova-3 | Amazon | transcription | — | — | — | — | — |
| nova-3-general | Amazon | transcription | — | — | — | — | — |
| nova-3-medical | Amazon | transcription | — | — | — | — | — |
| nova-general | Amazon | transcription | — | — | — | — | — |
| nova-phonecall | Amazon | transcription | — | — | — | — | — |
| playai-tts | — | speech | — | — | — | — | — |
| scribe-v1 | — | transcription | — | — | — | — | — |
| scribe-v2 | — | transcription | — | — | — | — | — |
| speech-02-hd | MiniMax | speech | — | — | — | — | — |
| speech-02-turbo | MiniMax | speech | — | — | — | — | — |
| speech-2-6-hd | MiniMax | speech | — | — | — | — | — |
| speech-2-6-turbo | MiniMax | speech | — | — | — | — | — |
| standard | — | speech | — | — | — | — | — |
| stt | — | transcription | — | — | — | — | — |
| stt-async-v4 | — | transcription | — | — | — | — | — |
| stt-async-v5 | — | transcription | — | — | — | — | — |
| tts | OpenAI | speech | — | — | — | — | — |
| TTS-1 | OpenAI | speech | — | — | — | — | 2023-11-06 |
| TTS-1 HD | OpenAI | speech | — | — | — | — | 2023-11-06 |
| tts-hd | OpenAI | speech | — | — | — | — | — |
| Voxtral Mini (latest) | Mistral | transcription | — | — | — | — | 2026-02-01 |
| Voxtral Mini TTS (latest) | Mistral | speech | — | — | — | — | 2026-03-01 |
| voxtral-mini-realtime | Mistral | transcription | — | — | — | — | — |
| voxtral-mini-transcribe-realtime | Mistral | transcription | — | — | — | — | — |
| whisper | OpenAI | transcription | — | — | — | — | — |
| Whisper | OpenAI | transcription | — | — | — | — | 2022-09-21 |
| Whisper 3 Large | OpenAI | transcription | — | — | $0.0023 | $0.0023 | 2024-10-01 |
| Whisper Large v3 Turbo | OpenAI | transcription | — | — | $0.0023 | $0.0023 | 2024-10-01 |
| whisper-base | OpenAI | transcription | — | — | — | — | — |
| whisper-medium | OpenAI | transcription | — | — | — | — | — |
| whisper-small | OpenAI | transcription | — | — | — | — | — |
| whisper-tiny | OpenAI | transcription | — | — | — | — | — |
There is no quality, Elo or rank column on this page. Every benchmark this site holds — all 17 of them, across 4 sources — scores text. None of them scores an image, a video or a voice, so any ranking here would be derived from price or release date and would measure nothing.
Common questions
- Which audio model is best?
- This site cannot answer that. Every benchmark it holds measures text, so there is no quality score for these models and no ranking. What this page gives you is the price each host lists, which is a fact.
- Why are the price units different from the rest of the site?
- Because the products are billed differently. An image model is usually priced per image, a speech model per second of audio, and some are priced per token like a language model. The column header names the unit. Do not compare across units.
- Where do these prices come from?
- The same daily catalog as every other price on this site: models.dev, LiteLLM and OpenRouter. A price belongs to a model and a host. Open the model page to see every host that lists it.
- Why are there fewer models here than on an image leaderboard elsewhere?
- The catalog covers models that a public API catalog lists with a price. A site that runs its own arena adds models it tests directly, including ones with no public API price. That is a different collection method, not a longer version of this one.