← Back to context

Comment by pembrook

7 hours ago

I see these conspiratorial arguments all the time and I think people massively overestimate the value of the average users tokens.

The problems with frontier models (design taste, ability to solve novel/difficult problems, etc) cannot be solved by throwing more slop from the average user at it.

Actually, most of the main deficiencies in current models stem from the fact that their data sets aren’t curated and specialized enough.

I don't think the goal of this data is necessarily model improvement.

I think it's marketing, advertising, and product refinement.

Ex: all the things Google wants your search data for.

It's somewhat silly to think the value of that data has changed much. Advertisers want to know what's popular and getting clicks and attention. Competitors want to know what features are getting used in their markets.

In the simplest case, think of this data as improving the harness, not the model.

  • I wonder how much less useful it is if I use those models for open code or similar. What are you really learning about me, other than the fact that I am a technical person, which you could know by the fact that I signed up for open router to start with.

The prompts contain sensitive personal data.

That would be valuable to advertizers for example.