← Back to context

Comment by rwz

4 hours ago

0.5t/s is still too slow even for email. For a moderately large inquiry (1MTok output, let's ignore the 4k context window limitation for now) it'll take the model around 23 days or uninterrupted execution to answer a single email.

Real world inquiries are gonna be much slower of course, but this setup is still too slow to do anything meaningfully useful I think.

Sure, you cannot do 1M, but there are plenty of useful questions you could ask that are far more modest. Simple Q+A, look at this function, how would you design X? All of those could have modestly sized outputs that would finish within a day.

I’m sure it’s possible, but I really struggle to think of an example that would result in a 1m token output.

  • it would be 1M tokens worth of work, with some small report for the end