Comment by gpm

1 day ago

All the alignment issues seem likely to be solved as soon as LLMs can actually do vision well IMO.

All modern multimodal models can do vision sufficiently well for web design, with the exception of respecting negative space and seeing poor padding/margins on text. The issue is that prompts are often egregiously underspecified so the design -> vision loop doesn't know how to refine.

  • There shouldn't need to be any prompting to result in correct alignment - like icons being centred. Under-specified prompts can justify the stylistic complaints where all the AI sites favour the same few designs, but not the simply erroneous positioning of things.

    I couldn't swear the issue is the vision and not a failure to choose to correct the issues after seeing them - but from interacting with them in other areas I suspect the former not the latter.