Comment by StilesCrisis

6 days ago

Read to the end: they aren't actually looking for the best-compressing output, because this quickly devolves into aaaaaaaaaaa. They keep a sliding window over a small portion of recent text and use that.

Basically I think the entire premise falls apart due to that choice--they forced an interesting-looking outcome by adjusting the algorithm until gzip started picking random slabs of letters instead of ever-larger repeating runs.

I had actually thought of doing this, but didn't for this exact reason. I knew I would have to fudge things to make it anything interesting.

Meh. That’s nothing compared to the amount of curation and tuning the LLMs are coerced with.

  • No--even a really tiny, underpowered model like GPT-2 with no system prompt produces coherent (though not necessarily desirable or correct) responses.

    • That's still an algorithm that underwent immense amounts of tuning. The tuning here on gzip is very simple, very few parameters, and generic. There is no reason to reject it.

      Also I don't know about calling GPT-2 "really tiny". You can get coherent responses out of 5-10M parameters.

      2 replies →