Comment by TuringNYC
1 day ago
Congrats on your app and love it so far! Already sent it to over a dozen family members. Curious about a couple of things
- I see only two employees on LinkedIn -- how were you able to QA all these different languages with just two people?!
- I tried Urdu and the app did quite well. But curious why you have two female voices and not any male voice?
- I realize Sesame is a much bigger team, but curious what you think they are doing that makes their voices feel so real and seamless. I dont think they do multiple languages so I think you have a harder problem of course.
Thank you so much for that!
We focused on testing and tweaking the most popular ones, we have not tested some of the niche ones. We have removed languages that users have told us have major issues, but there are still some left.
The voices are due to the quality of the TTS services that we use. Openi, 11labs, minimax. Some services don't have many or even 1 good voice. We will add more over time
Sesame also passes in the users voice into the TTS model so that it can vibe well with the users tone and mood, whereas we are just using raw TTS. Their latency is also very low, but this is not quite suitable for language learning.
In the future we hope to move to full voice to voice models, once those become mature and intelligent enough.