modelbenchmark.io

Audio models

Speech model prices

Transcription, speech synthesis and realtime voice models, with the price each host lists. No quality ranking: every benchmark on this site measures text. Language model prices.

Mode
WhisperOpenAItranscriptionfree$0.003free2023-09-01
GPT-4o Mini Transcribe (OpenAI)OpenAItranscription$1.25$1.25$5.002025-03-20
Gemini 3.5 TranscribeGoogletranscription$2$2.00$12.002026-08-26
GPT-4o Transcribe (OpenAI)OpenAItranscription$2.5$2.50$10.002025-03-20
gpt-4o-transcribe-diarizeOpenAItranscription$2.5$2.50$10.00
Gemini 3.1 Flash Live PreviewGooglerealtime$3$0.75$4.502026-03-26
gemini-2-5-flash-native-audioGooglerealtime$3$0.50$2.00
gemini-2-5-flash-native-audio-preview-09Googlerealtime$3$0.50$2.00
gemini-2-5-flash-native-audio-preview-12Googlerealtime$3$0.50$2.00
gemini-live-2-5-flash-native-audioGooglerealtime$3$0.50$2.00
gemini-live-2-5-flash-preview-native-audio-09Googlerealtime$3$0.50$2.00
nova-2-sonicAmazonrealtime$3$0.33$2.75
nova-sonicAmazonrealtime$3.4$0.06$0.24
Gemini 3.5 Live Translate PreviewGooglerealtime$3.5$3.50$21.002026-06-09
Gemini 3.5 Transcribe LiveGoogletranscription$3.5$3.50$21.002026-08-26
GPT-Realtime miniOpenAIrealtime$10$0.60$2.402025-10-10
gpt-realtime-2-1-miniOpenAIrealtime$10$0.60$2.40
gpt-4o-mini-realtimeOpenAIrealtime$11$0.66$2.64
gpt-realtimeOpenAIrealtime$32$4.00$16.00
GPT-Realtime-1.5OpenAIrealtime$32$4.00$16.002026-02-23
gpt-realtime-2OpenAIrealtime$32$4.00$24.002026-05-07
gpt-realtime-2.1OpenAIrealtime$32$4.00$24.002026-07-06
gpt-4o-realtimeOpenAIrealtime$110$5.50$22.00
asrtranscriptionfreefree
asr-largetranscriptionfreefree
basetranscription
base-conversationalaitranscription
base-financetranscription
base-generaltranscription
base-meetingtranscription
base-phonecalltranscription
base-videotranscription
base-voicemailtranscription
besttranscription
Canopy Labs Orpheus Arabic Saudispeech2025-12-16
Canopy Labs Orpheus V1 Englishspeech2025-12-19
chirpGooglespeech
chirp-3Googletranscription
eleven-multilingual-v2speech
eleven-v3speech
enhancedtranscription
enhanced-financetranscription
enhanced-generaltranscription
enhanced-meetingtranscription
enhanced-phonecalltranscription
Gemini 2.5 Flash Preview TTSGooglespeech$0.50$10.002025-05-01
Gemini 3.1 Flash TTS PreviewGooglespeechfreefree2026-04-15
generativespeech
gpt-4o-mini-ttsOpenAIspeech$2.50$10.00
gpt-live-1OpenAIrealtime
gpt-live-transcribeOpenAItranscription
gpt-realtime-translateOpenAIrealtime
gpt-realtime-whisperOpenAItranscription2026-05-07
gpt-transcribeOpenAItranscription
long-formspeech
Lyria 3 Clip PreviewGooglechat$0.040freefree2026-03-25
Lyria 3 Pro PreviewGooglechat$0.080freefree2026-03-25
lyria-002Googlespeech$0.060
muse-voice-transcribe-1-0Metatranscription
nanotranscription
neuralspeech
novaAmazontranscription
nova-2Amazontranscription
nova-2-atcAmazontranscription
nova-2-automotiveAmazontranscription
nova-2-conversationalaiAmazontranscription
nova-2-drivethruAmazontranscription
nova-2-financeAmazontranscription
nova-2-generalAmazontranscription
nova-2-meetingAmazontranscription
nova-2-phonecallAmazontranscription
nova-2-videoAmazontranscription
nova-2-voicemailAmazontranscription
nova-3Amazontranscription
nova-3-generalAmazontranscription
nova-3-medicalAmazontranscription
nova-generalAmazontranscription
nova-phonecallAmazontranscription
playai-ttsspeech
scribe-v1transcription
scribe-v2transcription
speech-02-hdMiniMaxspeech
speech-02-turboMiniMaxspeech
speech-2-6-hdMiniMaxspeech
speech-2-6-turboMiniMaxspeech
standardspeech
stttranscription
stt-async-v4transcription
stt-async-v5transcription
ttsOpenAIspeech
TTS-1OpenAIspeech2023-11-06
TTS-1 HDOpenAIspeech2023-11-06
tts-hdOpenAIspeech
Voxtral Mini (latest)Mistraltranscription2026-02-01
Voxtral Mini TTS (latest)Mistralspeech2026-03-01
voxtral-mini-realtimeMistraltranscription
voxtral-mini-transcribe-realtimeMistraltranscription
whisperOpenAItranscription
WhisperOpenAItranscription2022-09-21
Whisper 3 LargeOpenAItranscription$0.0023$0.00232024-10-01
Whisper Large v3 TurboOpenAItranscription$0.0023$0.00232024-10-01
whisper-baseOpenAItranscription
whisper-mediumOpenAItranscription
whisper-smallOpenAItranscription
whisper-tinyOpenAItranscription

There is no quality, Elo or rank column on this page. Every benchmark this site holds — all 17 of them, across 4 sources — scores text. None of them scores an image, a video or a voice, so any ranking here would be derived from price or release date and would measure nothing.

Common questions

Which audio model is best?
This site cannot answer that. Every benchmark it holds measures text, so there is no quality score for these models and no ranking. What this page gives you is the price each host lists, which is a fact.
Why are the price units different from the rest of the site?
Because the products are billed differently. An image model is usually priced per image, a speech model per second of audio, and some are priced per token like a language model. The column header names the unit. Do not compare across units.
Where do these prices come from?
The same daily catalog as every other price on this site: models.dev, LiteLLM and OpenRouter. A price belongs to a model and a host. Open the model page to see every host that lists it.
Why are there fewer models here than on an image leaderboard elsewhere?
The catalog covers models that a public API catalog lists with a price. A site that runs its own arena adds models it tests directly, including ones with no public API price. That is a different collection method, not a longer version of this one.