← Back to context

Comment by aunty_helen

1 day ago

Or, apples just so bad at this they’re fumbling the bag. Billions in cash on hand each quarter but don’t have the balls that zuck has to pay unreasonable money. They have their own hardware like google does but are talking about perplexity??? They have all data but can’t seem to get an llm that can set an alarm and be a chatbot at the same time?

Sometimes company’s just don’t do good enough.

> "They have all data but can’t seem to get an llm that can set an alarm and be a chatbot at the same time?"

This is actually one of the hardest frontier problems. The "general purpose" assistant is one of the singular hardest technical problems with LLMs (or any kind of NLP).

I think people are easily snowed by LLMs' apparent linguistic fluency that they impute that to capability. This cannot be further from the truth.

In reality a LLM presented with a vast array of tools has extremely poor reliability, so if you want a thing that can order delivery and remember your shopping list and remind you of your flight and play music you're radically exceeding the capabilities of current models. There's a reason successful (anything that isn't demoware/vaporware) uses of agentic LLMs tend to narrow-domain use cases.

There's a reason Google hasn't done it either, and indeed nor has anyone else: neither Anthropic nor OpenAI have a general purpose assistant (defined as being able to execute an indefinite number of arbitrary tools to do things for you, as opposed to merely converse with you).

  • You split up the tasks into sub agents. This is something my company builds on top of langgraph.

    • Sure, go try it and evaluate it rigorously end-to-end, over a sufficient number and variety of tools.

      For the purposes of the exercise, let's conservatively say, maybe ~2000 tools covering ~100 major verticals of use cases. Even that may be too narrow for a true general purpose assistant, but it's at least a good start. You can slice the sub-agents however you'd like.

      If you can get recall, for real user utterances (not contrived eval utterances authored by your devs and MLEs), over 70% across all the verticals/use cases/tool uses, I'd be extremely impressed. Heck, my thoughts on this won't matter - if you can get the recall for such a system over the bar you'd have cracked something nobody else has and should actively try to sell it to Google for nine figures.

      1 reply →

> Billions in cash on hand each quarter but don’t have the balls that zuck has to pay unreasonable money

It remains to be seen whether this was a smart move, or just flailing money at the wall

  • The difference is it’s a move. Actually doing something rather than putting out internal PR.

    Zuck tried and flailed with the metaverse. That was a huge waste, but he can afford it and fortune favours the brave.

    • The Metaverse was a waste of billions of dollars to develop a product that nobody wanted. In no world was that a smart business move, or one that should be emulated. Doing nothing is better than flushing money down the toilet.

> They have all data but can’t seem to get an llm that can set an alarm and be a chatbot at the same time?

This does seem like an embarrassing fail, but even Google has not completed replacing Assistant with Gemini. There have also been lost functionality (maybe temporary) in the process.

they are not talking about perplexity; the endless rumor mill talks about perplexity. The same that has them buying everything from Disney to Porsche to Nike for decades.

Undercut the competitors by charging less. Apple can afford to run its product at a loss.