Comment by redox99
2 days ago
(mumbles) Shannon entropy... Kolmogorov complexity... something, something...
On a more serious note, it depends on your cutoff for "entire human knowledge". It's easy to prove for a generous interpretations of "entire human knowledge" that it can't be done, but hard for something like "all useful human knowledge".
Wikipedia (text) is about 25gigs compressed. I think that's a reasonable starting point.
IDK, 30GB is a lot of data when we're talking about text!
Moby Dick, uncompressed, is ~1MB. Compressed, it's around 500KB.
I feel fairly certain that one could fit all of the textual knowledge required to cultivate a world-class <insert name of preferred professional knowledge worker> in <60,000 Moby Dicks. (Arguably in <5,000 Moby Dicks with intense effort/pruning).
I think a specialized model could squeeze all you need to know about a certain profession in 30GB. But not all professions at once, which is what these models try to do.
Or maybe not, maybe there's a world model needed for human level at any profession that is very hard to quantify and requires more than 30GB by itself.
Yea to be clear I think >70% of the information is not profession specific.
I just think about all the content I’ve consumed in my life to become a professional software developer and I would be very surprised if it couldn’t be adequately represented by <30GB of uncompressed text. Most of the work was in “training”, not really in data.
The “foundational overlap” of K-12 education is identical for most professions with 2-8 years of “finishing” on top.
My mental model is that the budget is spread across maybe 20% genetics (most of our instinctive/genetic information is surely pretty useless), 50% k-12 education, 30% for professionally-specific knowledge.
3 replies →