← Back to context

Comment by rdiddly

8 hours ago

It must have been used by humans before LLMs, because where do you think they learned it?

The training data set is a huge corpus of text from a wide variety of styles and time period. A human writer knows to adjust their writing style to fit their intended audience. A skilled LLM user knows to direct the LLM to generate in the appropriate style. A naive LLM user doesn't, so that's how we get this uncanny valley ick, where the text feels wrong.

All the marketing materials that got used for training. Marketing speak is not my cup of team.