Comment by CompoundEyes
7 hours ago
I do think it’s the wizard not the wand at this point given a decent model. These benchmarks don’t have the wizard.
Otherwise I wouldn’t see others in the exact same codebase struggle and underutilize agents while others thrive using the exact same ones.
In other words, we're still in the era of centaur chess.