Comment by OtherShrezzing
10 hours ago
This page is (somewhat ironically) so extremely laden with Claude-speak that it's difficult to find the information in all the noise. But once you've waded through everything, you see these facts:
>What the test measures: A model is given a passage and a fixed set of questions with short, checkable answers — a date, a name, a count.
So, a model is given content which is especially amenable to compression, and asked to reproduce it under certain constraints, like...
>Why isn’t the plaintext baseline 100%? Answering questions about an uncompressed passage in plaintext scores ~91%.... a correct answer worded differently scores as a [failure]
Models can (and do) give objectively correct answers, but are penalised for not having some kind of omniscient knowledge of the implementer's phrasing preferences.
If this phenomenon is emergent in models, this benchmark is not proof of it in any meaningful way.
For me the really interesting part was:
> Haven’t we seen LLMs do this already?
> Yes, BabelTele (arXiv, June 2026) demonstrated that LLMs can encode text in compact, non-standard forms — omnilingual word fragments, symbols, emoji — that other models recover with high fidelity (99.5% semantic fidelity at 27.9% of original length, by their metrics), including cross-model transfer, agent memory, and multi-agent communication. It proves the general phenomenon: human readability is not a requirement for model-to-model text.
I remember people testing early GPT-4 (2023?) in similar ways, to compress text, it would emit a string of strange text, Unicode, emojis, but was able to decode the compressed version very reliably.
This seems to cut usage by another ~50%, at the cost of being incomprehensible to humans.
Once, a GAN model that was trained to convert between satellite images and drawn maps was caught encoding the original satellite image in imperceptible dots
The field of ML is Goodhart's law reified. We might have temporarily forgotten some of the basics of the field amidst this LLM craze.
https://arxiv.org/abs/1712.02950 (2017)
…if you’re curious and missed that one like I did. Snack-sized paper with lots of satisfying visual examples.
Yeah! LLM compression is a spectrum between 'normal stuff we can read' and vectors. Depending on trust in the models and the desire for compression, there's a choice to be made on which formats you want to allow. Pretty interesting stuff.
Good for chain of thought, perhaps?
[dead]
A useful critique, thanks.
I would say that the grader has the same threshold for whatever answer it receives, and is equally harsh on whichever it grades. Any scores above the baseline (1.0) are really claims about parity, rather than better understanding in the compressed format.
The decoder step is a model expanding the cablese to regular text, not having seen the initial question. A separate model instance then reads that regular text and answers. And the result is still at parity with the plaintext record.
Had cablese knocked out information, that wouldn't have been the result, would it?
Ha, even you quoting the Claude-ese made me zone out of your comment and switch tabs off hacker news. As soon as I saw "What the test measures".
I only realized why I'd switched tab after I'd done it!