Comment by bonoboTP

9 hours ago

Text is not the only input. You can provide 3d block outs with rudimentary animation, annotated images with arrows etc, voice recordings of one person acting out some emotion then mapping that to a different character's voice, other uses of video to video, etc.

There could easily be at least some time period of low skilled ugly people acting in approximate but shitty ways in cheap sets just to give an input reference to a model and then describing the differences in text, yielding gorgeous people speaking with prestigious accents doing stuff in fancy locations in the output.