← The work indexApplied machine learning
Training & evaluation / 2026
A voice for long-form narration
Fine-tuning XTTS v2, then listening closely to what changed.
Illustrated study / A voice for long-form narrationExperiment
A technical study in improving Bible narration through dataset expansion, checkpoint selection, and inference tuning. The article includes audio comparisons and the measurements behind the choices.
More training wasn’t the answer.
The best validation checkpoint occurred at step 10,725. Training continued to step 35,750, but evaluation loss increased. Selecting the right checkpoint mattered more than using the last one.
Measure, then listen.
I compared pacing, pitch behavior, spectral similarity, and audible delivery. The case study distinguishes training signals from listening observations and describes the limits of the local evaluation.
Built with
Python / PyTorch / XTTS v2 / Audio evaluation
Next in the index
Box.tools
Web tools