Comment by chabska

4 hours ago

The training data set is a huge corpus of text from a wide variety of styles and time period. A human writer knows to adjust their writing style to fit their intended audience. A skilled LLM user knows to direct the LLM to generate in the appropriate style. A naive LLM user doesn't, so that's how we get this uncanny valley ick, where the text feels wrong.