← Back to context

Comment by embedding-shape

6 hours ago

Kind of feels like this applies to every single model, from Astra to Qwen, they all eventually lose track of the plot unless you feed it some human's input that can steer them right every now and then. The only difference is how often you need to do so, and also how often you want to do so heavily influences how good quality the results will be.

you are not false, but there is still difference. its just that the better models are correct more of the time and will better validate its own steps. glm sometimes will understand the plan start implementing and then forget part of it and then say it finished. or then take a wrong turn somewhere and not correct. but they will all happily proclaim they are correct till you question it.