Comment by visarga

2 years ago

> Meanwhile OpenAI, Anthropics, trains on AI generated data to improve their models, and it works.

They got a secret ace in their pocket - chat logs created with human in the loop. Of course those might still have errors, but much fewer. They can infer from a human response if it was accepted or not.

I think OpenAI generates at least 1B sessions per month and 2 Trillion interactive tokens. Those can go into the LLM again for analysis and synthetic content generation, or for RLHF with the whole conversation as guidance. Having access to the following interactions can shed light on previous answers.

Even more, they can correlate chats across days, presumably humans try out LLM ideas in reality and return for iteration. That way LLMs indirectly get real world grounding.

They can't directly train on chat transcripts, because they contain private information and other things you don't want appearing in answers. I doubt they even look at them unless you press the thumbs down, in which case they probably use it in some indirect way.

They might try to look for trends or what questions are popular of course.

  • That's exactly what they are doing and what you agreed to ans why some use other models or running something locally.

This is likely one of the main reasons why they're offering ChatGPT for free and running ChatGPT Plus at a loss.

  • > main reasons why they're offering ChatGPT for free and running ChatGPT Plus at a loss.

    As opposed to what though? Its not like there a huge demand for these apps that they can charge money. They have no option but to give it away for free .