Comment by minimaxir
1 day ago
All modern multimodal models can do vision sufficiently well for web design, with the exception of respecting negative space and seeing poor padding/margins on text. The issue is that prompts are often egregiously underspecified so the design -> vision loop doesn't know how to refine.
There shouldn't need to be any prompting to result in correct alignment - like icons being centred. Under-specified prompts can justify the stylistic complaints where all the AI sites favour the same few designs, but not the simply erroneous positioning of things.
I couldn't swear the issue is the vision and not a failure to choose to correct the issues after seeing them - but from interacting with them in other areas I suspect the former not the latter.