Whisper fine-tuning
PyTorchHugging Face TransformersPythonHugging FaceSpeech RecognitionFine-TuningDeep LearningNatural Language Processing
The problem
How does adapting a small speech model to a specific language change recognition?
How I solved it
Fine-tuned the 37.8M-parameter Whisper-Tiny model on Common Voice 11.0. Configured FP16, gradient checkpointing, TensorBoard and best-checkpoint selection by WER. The V2 run used 8,000 training steps with evaluation and checkpoint saving every 1,000 steps.
- Fine-tuned Whisper-Tiny for French and German
- Created 2 model variants for each language
- Evaluated performance against base model
- Achieved improved recognition accuracy
The results
- relative WER reduction
- 25.62%
- German evaluation utterances
- 16,082
The V2 notebook records word error rate falling from 43.49% to 32.35% across 16,082 German utterances: 11.14 percentage points, or 25.62% relative reduction. Published German and French model variants with training logs.
Evaluation context. Stored V2 notebook inference results, not a fresh rerun. The separate V1 result of 31.43% and V2 model-card result of 32.3327% use different reported evaluation records.