← Back to context

Comment by GenerWork

4 hours ago

I'm not sure I understand why they would route it to other models. It can't be that they don't have enough compute. Maybe worried that the answer from their own models would be bad? Doesn't really make sense, but I could be missing something.

> It can't be that they don't have enough compute.

How do you know?

Every cloud provider (AWS, Azure etc) is struggling with meeting LLM demand.

Source: first hand info

i agree, and it's what made me spend time exploring today.

fair q. call_ ids and gAAAAA blobs arent damning but the rs_ reasoning ids embed a unix timestamp that matches the session to the second, then OpenAI's 819x marker. plus the summary is in OpenAI's summarizer voice.

its just a best guess.