Tts: Amount of data to train own voice

Created on 23 Apr 2019  路  2Comments  路  Source: mozilla/TTS

Hey, I am trying to build a model with your version of Tacotron and my own data in US english.

As I am collecting data and formatting it, I am wondering what is the amount of data necessary to start getting some good results for the target voice? Does anyone know any empirical experiments so that I can set a target.

I have around 3/4h right now.

Thanks for the hard work on the repo

Most helpful comment

Amazon has some papers about the amount of the data needed for TTS. I'd say you can try to finetune one of the released models with your own data. That'd be the easiest way to go. You can also start from scratch but I'd guess the data is not sufficient. In general, my personal estimation is around 15 hours for something reasonable.

All 2 comments

Amazon has some papers about the amount of the data needed for TTS. I'd say you can try to finetune one of the released models with your own data. That'd be the easiest way to go. You can also start from scratch but I'd guess the data is not sufficient. In general, my personal estimation is around 15 hours for something reasonable.

Alright thank you very much! Will try that, and tell you if that improves it 馃憤

Was this page helpful?
0 / 5 - 0 ratings

Related issues

joshuadeisenberg picture joshuadeisenberg  路  3Comments

erogol picture erogol  路  9Comments

vcjob picture vcjob  路  7Comments

WeberJulian picture WeberJulian  路  6Comments

PetrochukM picture PetrochukM  路  9Comments