Comment by dhosek

6 years ago

I think the Chinese text that comes out confuses Google translate. I took the whole first sentence of Hamlet's soliloquy which compressed to 䮛趁䌆뺜㞵蹧泔됛姞音逎贊 and plugged that into Google Translate. It came back with "Commendation." The reverse translation is 表彰

10 comments

dhosek

james412 6 years ago

It's not Chinese text, it's an arithmetic-coded stream of bits mapped so the bits fall within the range of some codepoints. It's basically a variant of base64 except for Unicode.

(Side note: aren't these codepoints very expensive to encode in UTF-8? It seems there must be a lower-valued range more suited to it)

toast0 6 years ago
The page for base32768 has some efficiency charts for different binary to text encodings on top of different UTF encodings, as well as how many bytes you can use them to stuff in a tweet. Depends on where you're going to house the data, I guess.
https://github.com/qntm/base32768
- infogulch 6 years ago
  
  In addition to being 94% efficient in UTF-16 (!), this reveals some additional reasons why one might want to optimize for number of characters: fitting as many bytes as possible into a tweet which is bounded in the number of characters not bytes.
greenshackle2 6 years ago
Yeah I don't understand why it's using CJK, the page claims:
> each compressed character holds 15 data bits by using the CJK and the Hangul Syllables unicode ranges.
In UTF-8 these characters take 3 bytes each. Which makes it less space efficient than base64 (60% overhead vs 33% overhead).
The CJK/Hangul scheme has more information per character but I'm not sure where that matters.
- jkhdigital 6 years ago
  
  It's so that the output uses printable characters, that's all. The raw output would actually just be random bits, or at least something approximating random bits if GPT-2 is as good as we hope it is.
  
  2 replies →
willcipriano 6 years ago

It's probably similar to this: https://pieroxy.net/blog/pages/lz-string/index.html
Check the 'How does it look?' section.

dheera 6 years ago

This is Chinese characters mixed with Korean characters and this is pretty much never done by humans. It is analogous to mixing English and Heiroglyphics and typing out some gibberish with both.

The author might as well have included the rest of the Unicode range including Arabic, Emoji, and math symbols.

dhosek 6 years ago

I hit the reverse button again, and got "recognition" which translated back to 承认 which finally got into a closed loop to recognition and back to the same Chinese text.