← Back to context

Comment by miguelspizza

7 hours ago

You can easily get under 100ms per decision in browser with a Qwen based decison model: https://alxnahas.github.io/strands-decider-web/?backend=engi... There is nothing about Qwen or Laya that make one or the other better at running in the browser. It is really just: how large is your model and how does KV cache expand as input length increases

I guess it depends on your application, but 100ms is pretty slow. Plus there is the model footprint (memory + cpu/gpu).