← Back to context

Comment by Wowfunhappy

2 days ago

> You just can't compress the entire human knowledge into a 30GB file.

...I'm just asking questions here... how sure are we of this?

If you'd asked me six years ago whether we could compress all of human knowledge into a 1 TB file, I would have said no, and yet here we are.

If you can do 1 TB, why not 30 GB? It's, like, within one order of magnitude.

Amusingly enough, someone from a frontier lab could probably answer this empirically. Years ago Microsoft demonstrated training LLMs on synthetic text-- books rewritten by LLMs to be more concise and more accurate. https://arxiv.org/abs/2306.11644 It is well known that Anthropic extensively uses synthetic text in training. You could probably get good 50tile, 90tile, 99 etc numbers just from the size of the training materials on Anthropic servers.

(mumbles) Shannon entropy... Kolmogorov complexity... something, something...

On a more serious note, it depends on your cutoff for "entire human knowledge". It's easy to prove for a generous interpretations of "entire human knowledge" that it can't be done, but hard for something like "all useful human knowledge".

  • IDK, 30GB is a lot of data when we're talking about text!

    Moby Dick, uncompressed, is ~1MB. Compressed, it's around 500KB.

    I feel fairly certain that one could fit all of the textual knowledge required to cultivate a world-class <insert name of preferred professional knowledge worker> in <60,000 Moby Dicks. (Arguably in <5,000 Moby Dicks with intense effort/pruning).

    • I think a specialized model could squeeze all you need to know about a certain profession in 30GB. But not all professions at once, which is what these models try to do.

      Or maybe not, maybe there's a world model needed for human level at any profession that is very hard to quantify and requires more than 30GB by itself.

      4 replies →