Comment by skolos
3 hours ago
Interesting that this is here. I used whistle (and bunch of other things) to take ownership of my echo show. It now doesn't dial to Amazon at all - it does all processing locally with its own CPU and connects to my homeassistant for home automation. My initial setup involved qwen asr (1.7b model) running on rtx 5080. Compared to that, whistle was really bad (out of 170 messages, qwen recognized correctly 168, whistle - 70), but I adjusted whistle to work like jev - instead of free form transcription it recognizes only select set of templates (I trained tiny network with 10,000 generated utterances to translate whistle final state to probabilities within templates). The precision went up to 164/170 - almost matching qwen. By the way - I'm speaking with heavy accent.
This feels like we're going back to the past again with CMU's Sphinx4 in Java. It worked way better than I would ever have expected it to for being more than a decade old. It relied on the user defining a grammar of valid words and different flows through a standard format (Java Speech API Grammar Format). I wonder if we'll approach that again for these models just like how MCPs act like WSDLs in spirit. Great results getting whistle working so well for you!
Just ran some tests on my phrases:
- my whistle setup (tested can run on echo show with <1s response): 206 correct out of 208
These were not tested on echo show yet - on my pc for now:
- Vosk small, phrase grammar: 184/208
- Speech-to-Phrase 1.4.3 (Kaldi): 118/208
- suggested sphinx: 63/208 and on some examples took 13s on i7 14700k pc
looks like my customized whistle works better for me than these alternatives. But more testing wouldn't hurt.
> By the way - I'm speaking with heavy accent.
I chuckled at this because my inner voice had an accent as I was reading your comment, due to your writing style.
Excuse me, could you write slower? I couldn't understand you.
Sorry, curious non-native speaker here. Which accent? And which telling patterns made you think of it?
Not the person you're responding to, but as soon as I read stuff like "I trained tiny network" I tend to imagine Russian. It's the missing "a". Same thing happens in "I'm speaking with heavy accent".
Missing articles in English is the usual give away for someone with a slavic native language such as Russian.
Have you done a blog or YouTube about this ?
It is all custom made and not very reproducible. I'm working on reproducible setup and once it is done will publish it here: https://blog.kvit.app
If you are purposely limiting yourself to select templates, even fairly complicated templates, then classic voice recognition is perfectly sufficient.
With a restricted grammar, built in Windows voice recognition, all on device, has managed this exact use case quite well for over a decade. I used it to try and build a clone of the various paid apps that allow you to issue orders to Arma soldiers with voice commands