Comment by kingstnap
20 hours ago
Yesterday night I was doing a project with QwenTTS 1.7B.
After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).
I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.
So yeah the cat is out of the bag for sure.
Yeah I was doing that at the beginning of this year with voice samples locally from hollywood-stars with Qwen3-TTS.
It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.
Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D
Did you make her breakfast?
Any chance you could share a bit more detail? I’d love to try this myself but could use some proven structure / approach.
Same, I'd love a link.
Its so crazy to me how prevalent bad AI voices are, when local models can do such good AI voices
Very few people explore the options they have and tend to stick with the first thing that works.
Arguably, no AI voice should sound like a human voice:
https://youtu.be/M-IVVJkZnuo?t=236
uninspired
so one of my interests is reducing the payload size for video games
the vast majority of the image sizes have been audio recordings, and its been that way in different qualities for the last two decades. this is still the case as more varied and comprehensive audio is pursued by studios at unfathomable expense and still failing to cross a bar of realism
good voice models are just a few gigabytes in comparison and can supplant all of that, and be run locally at this point. Future ubiquitous hardware configurations in consumer devices will make inference dedicated and computationally cheaper and faster
although AAA studios are hamstrung and will be deeply unpopular if they stopped booking voice actors
everyone else who would have never had the capital for voice actors will just use this and have richer experiences until they themselves are AAA studios from the market buying their rich experiences
this will vastly supplant the assumed and uninspired “tricking humans” use case from that video. once it crosses a threshold of ease, the applications will expand