← Back to context

Comment by gchamonlive

7 hours ago

There was this post a few days ago https://news.ycombinator.com/item?id=49797323

It had this to say in the linked post:

  This led to the natural question: can gzip do language modeling? (...). Here’s some real, unedited output after priming it on tiny Shakespeare:

  gzipt --corpus data/tinyshakespeare.txt --prompt $'MENENIUS:\n' --length 200

  MENENIUS:
  'Though all at once canq

  MARCIUS:
  Pray now, nocamest thou to a morsel.

  LARTIUS:
  Hence, and
  I' the end admire, where G
  again; and after it ag .

Now thinking back, what's missing so that gzip could unwind the correct body of work from Shakespeare is just a correct sequence of bytes. One way to arrive at this is by just getting the body of work and doing the inverse, compressing it to get that golden sequence of bytes.

The other is what thinking does, it tries to predict the missing sequence of tokens from a high entropy source, the prompt, in order to increase the likelihood of correctly decompressing the desired results from its weights.

How can things be compressed without losing information or structure?

Like for text, what would that involve? How do you compress a string or multi-line string without losing information and hopefully structure (paragraphs, would it be like replacing periods and the following space with just sticking the starting capitalized letter of the following word to the previous sentence's last letter and when it decompresses theres some kind of note that converts that back into the. First letter of the next sentence

  • Other than what the other mentioned about finding more efficient representations, you can also compress by pre-agreeing on some common terminology.

  •     How can things be compressed without losing information or structure?
    

    Because the initial content is rarely the most efficient representation, so it's possible to store fewer bytes that can deterministically be converted into the original.

        Like for text, what would that involve?
    

    Most compression algos don't care what information you're compressing. All they see (all they need to see) is bytes. It ends up being way more sophisticated than removing repeated periods and whitespace.

    Like if you had eight boxes of loose lego, simply shuffling around the boxes wouldn't give you much in the way of reducing the space the legos take up. but if you took the legos (bytes) themselves out of the boxes, you end up saving a lot more space.