Aiola/whisper-medusa-v1 was trained on the LibriSpeech da... | Aiola/whisper-medusa-v1 was trained on the LibriSpeech da...
Aiola/whisper-medusa-v1 was trained on the LibriSpeech dataset to perform audio translation.
Whisper Medusa
Whisper is an advanced encoder-decoder model for speech transcription and translation, processing audio through encoding and decoding stages. Given its large size and slow inference speed, various optimization strategies like Faster-Whisper and Speculative Decoding have been proposed to enhance performance. Our Medusa model builds on Whisper by predicting multiple tokens per iteration, which significantly improves speed with small degradation in WER. We train and evaluate our model on the LibriSpeech dataset, demonstrating speed improvements.

Training Details
aiola/whisper-medusa-v1 was trained on the LibriSpeech dataset to perform audio translation. The Medusa heads were optimized for English, so for optimal performance and speed improvements, please use English audio only.

Usage
To use whisper-medusa-v1 install whisper-medusa repo following the README instructions.https://huggingface.co/aiola/whisper-medusa-v1 aiola/whisper-medusa-v1 · Hugging Face