Comment by StilesCrisis
6 days ago
Read to the end: they aren't actually looking for the best-compressing output, because this quickly devolves into aaaaaaaaaaa. They keep a sliding window over a small portion of recent text and use that.
Basically I think the entire premise falls apart due to that choice--they forced an interesting-looking outcome by adjusting the algorithm until gzip started picking random slabs of letters instead of ever-larger repeating runs.
I had actually thought of doing this, but didn't for this exact reason. I knew I would have to fudge things to make it anything interesting.
Meh. That’s nothing compared to the amount of curation and tuning the LLMs are coerced with.
No--even a really tiny, underpowered model like GPT-2 with no system prompt produces coherent (though not necessarily desirable or correct) responses.
That's still an algorithm that underwent immense amounts of tuning. The tuning here on gzip is very simple, very few parameters, and generic. There is no reason to reject it.
Also I don't know about calling GPT-2 "really tiny". You can get coherent responses out of 5-10M parameters.
2 replies →