Comment by minimaxir

1 day ago

All modern multimodal models can do vision sufficiently well for web design, with the exception of respecting negative space and seeing poor padding/margins on text. The issue is that prompts are often egregiously underspecified so the design -> vision loop doesn't know how to refine.

There shouldn't need to be any prompting to result in correct alignment - like icons being centred. Under-specified prompts can justify the stylistic complaints where all the AI sites favour the same few designs, but not the simply erroneous positioning of things.

I couldn't swear the issue is the vision and not a failure to choose to correct the issues after seeing them - but from interacting with them in other areas I suspect the former not the latter.