← Back to context Comment by Tsarp 7 hours ago Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model 8 comments Tsarp Reply forgot-my-pw 6 hours ago It might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=co... rvz 7 hours ago [flagged] jcims 7 hours ago We're allowed to have our ceremonies. kridsdale3 6 hours ago Thank you. If this whole thing isn't fun, it isn't worth doing. user43928 7 hours ago You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for? TylerE 6 hours ago Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes. 2 replies →
forgot-my-pw 6 hours ago It might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=co...
rvz 7 hours ago [flagged] jcims 7 hours ago We're allowed to have our ceremonies. kridsdale3 6 hours ago Thank you. If this whole thing isn't fun, it isn't worth doing. user43928 7 hours ago You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for? TylerE 6 hours ago Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes. 2 replies →
jcims 7 hours ago We're allowed to have our ceremonies. kridsdale3 6 hours ago Thank you. If this whole thing isn't fun, it isn't worth doing.
user43928 7 hours ago You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for? TylerE 6 hours ago Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes. 2 replies →
TylerE 6 hours ago Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes. 2 replies →
It might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=co...
[flagged]
We're allowed to have our ceremonies.
Thank you. If this whole thing isn't fun, it isn't worth doing.
You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.
2 replies →