Hi all,
today, we have added another 20 hrs to Ukrainian data set. Also, we have added 53hrs of Polish.
Lastly, we have added 190hrs of French dataset, which should be available around 17:00h CEST on our server. Enjoy and have fun. French was a bit of a nightmare, but we think we have "conquered it" :)
All info and updates are on: http://www.m-ailabs.bayern/en/the-mailabs-speech-dataset/ website...
Nice dataset ! How did you do the alignment btw, by hand?
No, we have developed a toolset to do the alignment, but you need to prepare the data upfront accordingly. The preparation can take somewhere between 3-4 days (French: 15 days :-) per language. But then, running through the toolset takes about 2-3 hours + QA. Normally, we can do one language dataset of 100hrs in about 5-8 days. But French took significantly longer.
Also, we do this only if we have no other urgent projects going on :). Our goal is to create the largest free dataset with multiple languages for speech recognition/synthesis. Let's see...
@imdatsolak great work! Looking forward to the finished version for French! wavs sound awesome!
@imdatsolak Is it possible to use your toolset for alignment ? I would like to prepare data for Belarusian language.
M-AILABS data is really great, I wonder why the website is down at the moment, hope this will get fixed soon..
Most helpful comment
No, we have developed a toolset to do the alignment, but you need to prepare the data upfront accordingly. The preparation can take somewhere between 3-4 days (French: 15 days :-) per language. But then, running through the toolset takes about 2-3 hours + QA. Normally, we can do one language dataset of 100hrs in about 5-8 days. But French took significantly longer.
Also, we do this only if we have no other urgent projects going on :). Our goal is to create the largest free dataset with multiple languages for speech recognition/synthesis. Let's see...