Whisper: Robust Speech Recognition and Translation
Whisper is a family of automatic speech recognition and speech translation models from OpenAI, introduced in the paper Robust Speech Recognition via Large-Scale Weak Supervision. The flagship large-v3 checkpoint was trained on 1 million hours of weakly labeled audio and 4 million hours of pseudo-labeled audio, and delivers a 10 to 20 percent error reduction over large-v2.
On Vast, Whisper is served through the OpenAI-compatible audio transcription endpoint, so any client that already speaks the OpenAI transcription API can point at the instance without code changes.
Key Features
Multilingual Transcription
Whisper transcribes speech in 99 languages, with automatic language detection. Robustness to accents, background noise, and technical language comes from the scale and diversity of its training data, and it generalizes zero-shot to many domains without fine-tuning.
Speech Translation
Beyond same-language transcription, Whisper translates speech from any supported language directly into English text in a single pass.
Timestamps
The model produces sentence-level and word-level timestamps, making it a complete pipeline from raw audio to captions and searchable transcripts.
Drop-in API
Transcription requests use the same request and response shape as the OpenAI audio transcription API, including the choice of plain text, JSON, or subtitle-style output.
Use Cases
- Transcribing meetings, interviews, lectures, and podcasts
- Generating subtitles and captions for video content
- Translating foreign-language audio into English text
- Building voice-driven applications on a private transcription API
- Batch transcription of audio archives
Architecture
Whisper is a Transformer-based encoder-decoder model with 1.55 billion parameters in its large-v3 configuration. Audio is converted to a 128-bin Mel spectrogram, up from 80 bins in large-v2, and processed in 30-second windows, with long-form audio handled through sequential or chunked decoding. On the Open ASR Leaderboard, large-v3 reports a mean word error rate of 7.44 across evaluation datasets.