Comment by Tsarp 8 hours ago Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model 8 comments Tsarp Reply forgot-my-pw 7 hours ago It might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=co... rvz 8 hours ago [flagged] jcims 8 hours ago We're allowed to have our ceremonies. kridsdale3 7 hours ago Thank you. If this whole thing isn't fun, it isn't worth doing. user43928 8 hours ago You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for? TylerE 8 hours ago Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes. 2 replies →
forgot-my-pw 7 hours ago It might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=co...
rvz 8 hours ago [flagged] jcims 8 hours ago We're allowed to have our ceremonies. kridsdale3 7 hours ago Thank you. If this whole thing isn't fun, it isn't worth doing. user43928 8 hours ago You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for? TylerE 8 hours ago Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes. 2 replies →
jcims 8 hours ago We're allowed to have our ceremonies. kridsdale3 7 hours ago Thank you. If this whole thing isn't fun, it isn't worth doing.
user43928 8 hours ago You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for? TylerE 8 hours ago Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes. 2 replies →
TylerE 8 hours ago Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes. 2 replies →
It might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=co...
[flagged]
We're allowed to have our ceremonies.
Thank you. If this whole thing isn't fun, it isn't worth doing.
You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.
2 replies →