Comment by bnchrch

20 hours ago

I don't think the "don't really save you money" hot take holds water in every case.

Coding, maybe.

But for operationalized/repeatable tasks it definitely does.

For example I have a workflow that I was running in April that effectively would cost $30k in token spend for each full run.

However now, with GLM 5.3-flash, we've brought the cost down to $7k-9k with our evals showing we've had no loss in recall, precision etc..