Kokoro 82M: Small, Fast, Open Text-to-Speech
Kokoro is an open-weight text-to-speech model with just 82 million parameters, published at hexgrad/Kokoro-82M. Despite its size, it delivers speech quality comparable to models many times larger: entered into the community TTS arenas shortly after release, it climbed to the top of the rankings while remaining fast and cheap enough to run interactively on modest hardware. The v1.0 release ships 54 voices across 8 languages.
Overview
Kokoro's design philosophy is efficiency. It builds on the StyleTTS 2 architecture with an ISTFTNet vocoder, using a decoder-only formulation with no diffusion process, so synthesis is a single fast forward pass rather than an iterative sampling loop. The entire model was trained for roughly a thousand A100 GPU-hours on a few hundred hours of permissively licensed and synthetic audio with IPA phoneme labels. Text is converted to phonemes by the companion misaki G2P library, and the model generates 24 kHz audio conditioned on a chosen voice embedding.
Key Features
- Tiny footprint - 82M parameters; loads in seconds and synthesizes far faster than real time on any modern GPU
- 54 voices, 8 languages - American and British English, Spanish, French, Hindi, Italian, Brazilian Portuguese, Japanese, and Mandarin Chinese, each with multiple named voices
- Voice blending - Voice embeddings can be interpolated, mixing two voices at an adjustable ratio to create new ones
- No diffusion - Decoder-only StyleTTS 2 + ISTFTNet design produces audio in one pass, keeping latency low and throughput high
- Truly open - Open weights trained only on permissive, non-copyrighted, and synthetic audio, with commercial use encouraged
Architecture
Kokoro pairs the StyleTTS 2 formulation with an ISTFTNet-based vocoder that reconstructs the waveform through inverse short-time Fourier transform layers rather than a heavy neural upsampling stack. A style vector selected from the voice pack conditions prosody and timbre, while the phoneme sequence from misaki drives content. Because there is no diffusion sampler and no large acoustic encoder at inference time, generation cost scales gently with text length, which is what lets an 82M-parameter model serve production traffic efficiently.
Use Cases
- Voiceover and narration: turn scripts, articles, or documentation into natural-sounding audio
- Application and game voices: low-latency speech for assistants, NPCs, and accessibility features
- Audiobook and podcast production: long-form synthesis where per-character cost matters
- Multilingual content: one model covering eight languages with consistent quality
- Custom voice design: blend existing voices to produce a distinct brand voice without training
On This Template
The template runs Kokoro inside ComfyUI through a dedicated TTS node pack. A ready-made workflow is preloaded: type or paste text, pick from the built-in voice list, adjust the speaking speed, and queue - the synthesized speech is saved as a FLAC file you can play directly from the output panel. The node also exposes voice blending, mixing any two voices at a chosen ratio. The template ships 41 voices across 7 languages: American and British English, Spanish, French, Hindi, Italian, and Brazilian Portuguese. The Japanese and Mandarin voices are not included, because their text-processing libraries are not installed. The model weights and voice packs are downloaded during provisioning, so the first generation starts immediately. The same workflow is also available through the included API wrapper for programmatic and serverless use.