Model Library/Whisper large-v3

Whisper large-v3

STT
Audio

Whisper transcribes and translates speech in 99 languages, served through the OpenAI-compatible audio transcription endpoint

On-Demand Dedicated 1xRTX 4090

Details

Modalities

audio

Recommended Hardware

1xRTX 4090

Estimated Price

Loading...

Provider

OpenAI

Family

Whisper

Parameters

2B

Context

448 tokens

License

Apache 2.0

Whisper: Robust Speech Recognition and Translation

Whisper is a family of automatic speech recognition and speech translation models from OpenAI, introduced in the paper Robust Speech Recognition via Large-Scale Weak Supervision. The flagship large-v3 checkpoint was trained on 1 million hours of weakly labeled audio and 4 million hours of pseudo-labeled audio, and delivers a 10 to 20 percent error reduction over large-v2.

On Vast, Whisper is served through the OpenAI-compatible audio transcription endpoint, so any client that already speaks the OpenAI transcription API can point at the instance without code changes.

Key Features

Multilingual Transcription Whisper transcribes speech in 99 languages, with automatic language detection. Robustness to accents, background noise, and technical language comes from the scale and diversity of its training data, and it generalizes zero-shot to many domains without fine-tuning.

Speech Translation Beyond same-language transcription, Whisper translates speech from any supported language directly into English text in a single pass.

Timestamps The model produces sentence-level and word-level timestamps, making it a complete pipeline from raw audio to captions and searchable transcripts.

Drop-in API Transcription requests use the same request and response shape as the OpenAI audio transcription API, including the choice of plain text, JSON, or subtitle-style output.

Use Cases

  • Transcribing meetings, interviews, lectures, and podcasts
  • Generating subtitles and captions for video content
  • Translating foreign-language audio into English text
  • Building voice-driven applications on a private transcription API
  • Batch transcription of audio archives

Architecture

Whisper is a Transformer-based encoder-decoder model with 1.55 billion parameters in its large-v3 configuration. Audio is converted to a 128-bin Mel spectrogram, up from 80 bins in large-v2, and processed in 30-second windows, with long-form audio handled through sequential or chunked decoding. On the Open ASR Leaderboard, large-v3 reports a mean word error rate of 7.44 across evaluation datasets.

Quick Start Guide

Choose a model and click 'Deploy' above to find available GPUs recommended for this model.

Rent your dedicated instance preconfigured with the model you've selected.

Start sending requests to your model instance and getting responses right now.