Comment by xp84
7 hours ago
This reminds me of how the thing iPhones had pre-Siri (so we're talking pre-2010), which was entirely offline, did a better job than even the most modern thing at "Play [one of the finite set of songs in my library]." I sometimes get absurd matches from bands I've never heard of, when the right answer is something right there in my library.
It's odd that the matching algorithm does not simply prioritize the music already in your library, but I see something similar in other domains where machine learning/information retrieval is used. E.g., in Apple Maps, I might have the map centered over my location and type in a restaurant nearby. Often, Apple Maps will find a restaurant with the same name on the other side of the country. This strikes me as an easy thing to fix (and Apple Maps has had this bug from the beginning), but if somebody knows something abou this, I'd love to know. Maybe it's harder than I imagine.
I know it makes me sound like a lunatic to anyone listening, but I always use the most condescending monotone with Siri in the car, like this. "Directions to Wal-Mart,... in... Beaverton,... Oregon." Doesn't matter if I've been to that Walmart 197 times since Apple Maps 1.0 came out, because if you don't be as explicit as , 1 out of 10 times, it'll decide that you must mean a random Walmart 12 states away. Or like "Wall Plastering Incorporated."
And even when it's getting it right, and if there's only one Walmart in Beaverton, it still to this day needs to ask "One option is Walmart on Expressway Road in Beaverton...." Maybe it's correct in its 0% confidence level there, since it's so bad, but... I don't get how you could design something that bad, even before LLMs existed. I feel like I could do better, even using their Speech-to-text engine, with the processing backend built of pure regexes and if/elses.
It's just useless for me now, the change happened some 5 or so years ago.
"Hey Siri, play [song]"
Leads to, take your pick:
- "You'll need to unlock your iPhone first."
- "I couldn't find [song] on Podcasts" (??????)
- "Playing [a totally different song]"
- "I couldn't find any music by [song, but it thinks it's a band]"
- "Playing music by [song, again it thinks it's a band]"
Yes!! I get all of those. My favorite is that 10% of the time it asks me on what app I want to play the music (Oh, Apple, you're suddenly deeply respectful of competing on an equal playing field?).
And don't forget whatever the current phrasing is for "I'm sorry, my shit's all fucked up" and "My network connectivity had a blip and I'm unwilling to retry" and "Even though I have on-device STT models, and now LLMs too, and an on-device database of your music, which is downloaded, I won't bother without the cloud.
Settings -> Accessibility -> Side Button -> under "Press and Hold to Speak" choose "Classic Voice Control"
This is because, in the case of a restricted set of possibilities, voice recognition circa 2000 was actually very very good.
If you can do something with an extremely limited vocab, voice recognition was fine using off the shelf microchips in the 70s, where you wired in a microphone connection and had discrete pins for output actions.
LLMs are basically only useful for utterly free form transcription, but that doesn't actually help you turn that into tasks to perform and parameters for those tasks
The core "problem" in voice recognition is that freeform speech is an abysmal UX paradigm and provides zero discoverability, and LLMs IMO have not improved the situation of actually doing anything with the resulting text.
The other day I tried to prompt Gemini 3 times to tell me what the heck the business with a weird sign I saw was. The first prompt worked with a stale location context and therefore was way off, the second prompt had to reach out to google servers, and came back with recognizing the physical space I was discussing, but told me that I was talking about an event that takes place in the museum next door that I had told the model was next door to the business in question, the third try it still seemed to understand where I was referencing, but insisted I couldn't possibly be talking about anything there.
It took 1 second on google maps to find exactly what I was referring to, which was the business in Google's system located at the exact map location the model had found.
I'm sick and tired of people turning to LLM and "AI" tools to pretend they are better, when the problem is that these companies don't even use existing good solutions because they just don't care.