A Spectre Is Haunting Unicode (2018)

4 years ago (dampfkraft.com)

In the 90s I worked on a project to digitize land registration in Taiwan.

In order to record deeds and property transfers, we needed to enter people's names and official registered addresses into the computer system. The problem was that some people used non-traditional writing variants for their names, and some of their birthplaces were tiny places in China with weird names.

Someone might write their name with a two-dot water radical instead of three-dot radical. We would print it out in the normal font, and the people would lose their minds, saying that it was wrong. Chinese people can be superstitious about the number of strokes in their name, so adding a stroke might make it unlucky, so they would not buy the property.

The customer went to the agency responsible for managing the big character set, https://en.wikipedia.org/wiki/CNS_11643 Despite having more characters than anything else on earth, it didn't have those variants. The agency said they would not encode them, because they were not real characters, just printing differences.

The solution was for the staff in the office to use a "font maker" program to create a custom font with these characters. Then they could print out the deeds using a Chinese variant of Adobe Acrobat, and everyone was happy.

  • That's a great story. The inability to represent a name with standard characters reminds me of when Prince changed his name to a symbol and they had to send all of the media floppy disks containing a custom font with a single character.

    https://nymag.com/intelligencer/2016/04/princes-legendary-fl...

    • Are you acquainted with Freur (which means, "Underworld 0.5" - Rick Smith and Karl Hyde in the '80s)?

      "Freur", or, "The squiggle we chose as the name for a band but that CBS Records insisted should at least have a pronunciation".

      I see it is not in Unicode (well, you can never really know if you do not try), nor I can find pieces to reconstruct it.

      The "freur" in foreground: https://d4q8jbdc3dbnf.cloudfront.net/user/6885/edb290c6183ac...

  • I've been told that this is also an issue in Japan, except the reason might more often be a matter of pride than superstition. It is supposedly one reason (of a few) why fax machines are still in common use in Japan.

    Later versions of Unicode support "Variation Forms" of Han characters as a way to be able to encode different variations. They are encoded as a Variation Selector code (U+E01000 and up) after the Han character. The forms are listed separate from Unicode versions in the "Ideographic Variation Database" <https://www.unicode.org/ivd/>. So far, it contains characters from a couple of Japanese dictionaries, a Korean and one from Macao/Hong Kong.

    • I knew someone who added an accent character to their name because everyone pronounced it wrong. She met someone bilingual who shot back that if she wants it pronounced that way she needs to add an aigue. So she did, and everyone still pronounced her name wrong.

      In fact going any place with her very nearly became an “are we living in a simulation” crisis for me because the number of times she would say her name and the other person would say it back incorrectly was… upsetting. The degree to which some people butchered her name, especially combining half of her first and last name into a completely different name, made us joke about buggy NPCs.

      I could imagine how in some cultures writing it incorrectly hurts as much as pronouncing it incorrectly. Or possibly moreso in places where multiple plausible pronunciations have to be negotiated via an introduction, which is the case in China, is it not?

      14 replies →

  • Forgot which country (iran, turkey..) but one diacritic on a phone text got a girl killed because it altered the meaning one word. Turning the sentence from loving to threatening or insulting.

  • That sounds equally fascinating, and a little madding.

    • Yep, and with pictographic writing systems it's a lot more common than latin... but even here we have X Æ A-12 Musk, and Prince's name symbol.

      Heck, my initials are totally non-standard.

  • > Chinese people can be superstitious about the number of strokes in their name, so adding a stroke might make it unlucky

    Why am I not surprised in the slightest?

This might be interesting read to those unfamiliar with CJK, but character bloat(?) isn't remotely a recent thing. It's actually at least a couple hundred years old.

The Kangxi dictionary (1716), an authoritative dictionary of Chinese characters, contains definitions for 47035 characters, even though only a couple thousand are in common use. Quoting from Wikipedia: "The dictionary was the largest of the traditional dictionaries, containing 47,035 characters. Some 40% of them are graphic variants, however, while others are dead, archaic, or found only once. Fewer than a quarter of the characters it contains are now in common use."

All of these archaic (or even bogus in some cases) characters found in the dictionary are now part of the Unicode standard, of course :) The unihan database even has a field that shows the page number where the character appears in the Kangxi dictionary. If you're wondering why 65536 characters isn't enough for everyone, the junk in Kangxi dictionary is a significant contribution.

  • Unicode is a mistake that could only have happened in turn of the century America.

    It is the distilled essence of the idea that you need to be inclusive of everyone along with a fundamental ignorance of what anyone who isn't American does.

    The idea that Chinese characters are glyphs in the same sense of Latin characters can only have come from someone who has never written Chinese.

    It is as stupid as demanding a glyph point for each possible integral, e.g. https://quicklatex.com/cache3/8c/ql_9739884527bd893429657272... and https://quicklatex.com/cache3/18/ql_c51509950f58a52253c696a4....

    Solutions for English are not solutions for all languages. You can tell because the solutions that were natively invented by people who spoke those languages were _not_ unicode.

    • JIS, Big5, UHC, and GB all use a codepoint-to-character approach. You're right to point out that many aspects of "multilingual" support are written by people who do not know another language and so end up being hopelessly misguided but it's not really fair to say that Unicode invented this and thrust it upon the CJK world. Every pre-existing system of representing 漢字 had a codepoint table (which Unicode references in their description of each character).

      Han Unification was in my view problematic but was driven by technical limitations (then again, if Simplified Chinese characters had also been unified I suspect there would've been more pushback to come up with a better solution, but ultimately Japanese was stuck with being the only one making a major compromise on that front).

      I don't think a stroke based or combination system would've been better for many reasons: https://news.ycombinator.com/item?id=32102093. And if you don't trust Americans who at least tried to learn about the subject matter, how much do you trust any other programmer (who has no interest in other languages) to be able to handle a more complicated system for representing and rendering 漢字?

      5 replies →

    • Han Unification was invented by Chinese people (in Hong Kong iirc), not by Americans. And even before unification, the national standards also had one code point per kanji/hanzi rather than building them up out of radicals, as opposed to the way eg flag emoji are done.

      JIS did this in 1978 for instance.

      It does appear that computerizing CJK languages has made them very different from handwriting them; native Chinese speakers now constantly forget how to write hanzi. But they did this to themselves.

      1 reply →

  • I think 'character bloat' is simply inherent to the writing system when characters are written by hand (now that perhaps most written communication is digital people can't use characters that are not already supported)

    Anyone can invent characters whenever they want, and it's only a question of them sticking or not.

    I think this is also one of the reasons for the Chinese tendency to push for unification and uniformity.

    • When it’s character based instead of alphabet based, I think it’s the equivalent of coming up with a new word in English, which is basically what you’re describing.

      Sometimes it’s mashing two previously unrelated ‘words’ together (aka the tons of compound characters in Chinese), other times it’s coming up with something completely new.

      Same rules apply though, if it doesn’t add value worth the trouble (or get mandated by the powers that be), it’ll eventually just die out or be a curiosity.

      Also, to keep it tech related:

      RISC = English CISC/VLIW = Chinese?

      9 replies →

    • I'm not sure I understand. Most European languages go through cycles where letters are added when languages are mixed together followed by periods of redundant letters disappearing. Old English had something like 39 letters. 'th' used to have its own letter: thorn.

      I think character proliferation in CJK languages are a result of each word having its own character. The proliferation isn't fundamentally a proliferation of characters, it's a proliferation of words, which happens all the time in all languages. But only in certain languages does this proliferation of words result in additional characters being added to the language.

      1 reply →

  • >Fewer than a quarter of the characters it contains are now in common use

    12K characters in common use is equally impressing for me as a non-Asian.

    • It's actually way fewer than that IRL. Japan's official list of commonly used Kanji only has 2136 characters. Taiwan's list has 4808, and the PRC's list has 3500 "frequent" characters with another 3000 supplementary "common" ones. Digitization has made it even easier to use these characters without recognizing the actual form or how to write them.

      4 replies →

    • More like 12k characters currently in use at all. Common use characters are a much smaller set than that. (3k or so?)

  • Does Unicode really need to store Chinese words? Is it impossible to deconstruct the glyphs into strokes, each stroke effectively being a character?

    • The problem with that would be that every software must know the intricate rules about combining glyphs, and if they guess wrong, users get garbage characters.

      Considering that the majority of code is written by people who don't know Chinese characters, it would result in never-ending issues, pretty much everywhere.

      Korean actually has a two-way system in Unicode. Every conceivable character (= syllable) possible in modern Korean has its own codepoint, which allows most software to display them correctly: from their point of view, it's just another CJK character.

      On the other hand, there is a Unicode area containing Korean sub-blocks ("jamo") that were used historically. In theory, you can combine them and get some pretty funky archaic syllables. Almost no software renders them right.

      2 replies →

    • In addition to the problems mentions by yongjik, even with the current system, very little software is even aware that the same codepoint should be rendered differently in different languages (返す needs to be rendered differently in every CJK locale) which often results in websites and programs using Chinese fonts for Japanese text (even if you've configured your language as Japanese). Having stroke breakdowns would not make this situation better because there are multiple ways to render the same stroke description and there aren't really systematic rules for how to correctly represent the Japanese (or Taiwanese or Korean) version of a character -- it's generally for historical reasons. If you were to try to actually represent the characters faithfully (in an attempt to avoid making every country unhappy with the way you've butchered their language), many characters would become unusable for text searching because the same "character" (from the perspective of a CJK native) would have a completely different representation in a way that a computer could not be able to identify as being the same (even a character as simple as 言う would have this issue).

      I dread to think what an enormous mess would result if every character was represented as a build-it-yourself instruction manual rather than allowing font authors to correctly represent the characters. This is also ignoring that (depending on the font style), the apparent strokes for a character can change between fonts in the same language (this is because the computer font stroke style and the written font stroke style can be different) -- by putting stroke decisions in the encoding you're introducing a layering violation since fonts should be deciding how characters are styled, not encoding format committees.

      Also nobody in China, Japan, nor Korea would switch to an encoding system so incredibly inefficient that more strokes results in more bytes being necessary to store the character (they already compromised with having 3-byte UTF-8 characters when JIS, GB, and Big5 all only required 2 -- and Japan was basically forced to compromise on Han Unification). This would've resulted in the failure of Unicode's mission to be the One True Encoding Format.

    • In the early days of computers some character systems were stroke-based because that used less memory than a 32x32 bit map. A kilobit of ROM (one character) could cost $10.

      Currently stroke-based systems are used for calligraphic effect. You could generate new font types, e.g. bold., but controlling the shape of strokes.

      Stroke systems are important for teaching character writing because the drawing order is rigorously prescribed. Once you learn the first couple hundred, you can pretty much guess future characters. Wrong order characters often look bad and suggest a non-Chinese speaker mis-copied them. (e.g. some tattoos)

    • Unicode has support for this, in the Ideographic Description Characters block (https://en.m.wikipedia.org/wiki/Ideographic_Description_Char...). However, it’s purely descriptive, and not designed for rendering.

      There are somewhat more sophisticated systems which define both the rendering and stroke decomposition of characters (e.g. CDL: http://guide.wenlininstitute.org/wenlin4.3/Character_Descrip...). The general workaround for characters that aren’t on Unicode would be to use one of these stroke description systems to create the character, then render it to an image and insert it.

I thought this was going to be about something like the massive security problem of homoglyph attacks being currently deployed in stuff like phishing baked into the standard at first glance of the title, but this ghost character business is pretty interesting. Japanese literacy requires you to know 2-4 meanings per 2,136 kanji characters (something like 6000+ in total possible meanings between these characters) just to be able to pass a university level literacy test, it's a massive amount of complexity to get right. Even if you just need basic literacy it's still about a thousand less than that, and there's even more than these I mentioned for further literacy competence. Furthermore each of these characters look funny if not unreadable if you write them down using the wrong order of strokes. I can see how mistakes might have been made even by native speakers of that language. The two kana syllabiaries are there of course and mixed in with the kanji, but if everything was written in that you wouldn't be able to achieve the same amount of information density, which is probably part of the reason they never switched over (I understand before world war 2 or so, the more rounded hiragana was for women while the more sword stroke like katakana was for men).

The Latin alphabet being boring, I spent some time going through ancient alphabets included in Unicode.

It gets pretty trippy, pretty quick.

As in "We don't have a clear idea what this rune was for, or what it means, but we see it in documents and so added it to Unicode."

https://en.m.wikipedia.org/wiki/Runic_(Unicode_block)

  • My favorite Unicode glyph is Multiocular O (ꙮ). There is only one recorded usage, by a 15th century russian monk, who decided to use it in phrase “many-eyed seraphim” instead of two regular letters ‘o’. So of course it was added to Unicode.

    https://en.wikipedia.org/wiki/Multiocular_O

    • It gets better: this glyph is bugged. Somehow, the guy responsible for adding it to Unicode somehow got the number of eyes wrong. Per his description, Unicode fonts represent it with 7 eyes, but after getting called out on Twitter he realized the original manuscript shows 10 eyes.

      This bug will be fixed in Unicode 15.

      3 replies →

    • This one has sat in the back of my head for a long time. I wonder if any of the little ligatures or strange letter variations people write today could be preserved in the same way, or if the shorthand systems could be.

  • > As in "We don't have a clear idea what this rune was for, or what it means, but we see it in documents and so added it to Unicode."

    Documents? I had the strong impression that there are no documents written in runes. A rune we only know by its occurrence in documents would be far more interesting for the existence of a document than it would be for its own sake!

    Compare what the page about Anglo-Saxon runes says about the corpus:

    > The Old English and Old Frisian Runic Inscriptions database project at the Catholic University of Eichstätt-Ingolstadt, Germany aims at collecting the genuine corpus of Old English inscriptions containing more than two runes in its paper edition, while the electronic edition aims at including both genuine and doubtful inscriptions down to single-rune inscriptions.

    > The corpus of the paper edition encompasses about one hundred objects (including stone slabs, stone crosses, bones, rings, brooches, weapons, urns, a writing tablet, tweezers, a sun-dial,[clarification needed] comb, bracteates, caskets, a font, dishes, and graffiti). The database includes, in addition, 16 inscriptions containing a single rune, several runic coins, and 8 cases of dubious runic characters (runelike signs, possible Latin characters, weathered characters). Comprising fewer than 200 inscriptions, the corpus is slightly larger than that of Continental Elder Futhark (about 80 inscriptions, c. 400–700), but slightly smaller than that of the Scandinavian Elder Futhark (about 260 inscriptions, c. 200–800).

    So across every runic system we know, we have under 600 texts, all of those texts are short inscriptions, and even to reach that number of samples we need to include texts that we aren't even sure contain any runes.

    • > Documents? I had the strong impression that there are no documents written in runes.

      One of the original goals of Unicode was to be able to computerize every document. I still have some old linguistics books in which characters have been handwritten into typed or even typeset text. So these are the types of documents being referred to: academic papers.

      Some fancy books have photographs of ancient writing; I’m not sure if Unicode tries to encode such sources and I pretty much doubt it (how would you even know what to call the symbols? You touch on this in your comment). However often they are attached to treatises that order the characters in some way (I.e. index an alphabet) in which case the first case above would apply.

      In other words: thanks to some scholars who wrote down and ordered runic alphabets, you can now discuss runes with your friends and colleagues through email.

      14 replies →

Can we talk about the artwork used?

https://dl.ndl.go.jp/info:ndljp/pid/1312837?itemId=info%3And...

https://philamuseum.org/collection/object/84871

Googling for Tsukioka Yoshitoshi brings up so much SEO that it is hard to find information in English. If anyone knows anything about it, I'd be appreciative for a pointer about its content/subject!

> At this rate they'll presumably be with humanity forever. Ψ

So, that's a really interesting thought. Perhaps our solution to a permanent reminder of nuclear destruction[1] could be hidden inside a plane of Unicode.

[1] https://en.wikipedia.org/wiki/Long-term_nuclear_waste_warnin...

  • Maybe Unicode will feature the same kind of warnings one day.

    > This Unicode range is not a place of honor. No highly-esteemed symbol is registered here.

    > What was here represented cultural signs that were considered powerful in our time.

If you type 彁 on google translate and set it to detect language it will switch to Chinese and translate it to "lingering". If you switch to Japanese no translation will happen.

Also if you google search for 彁 one of the results will be this video [!!!!seizure warning!!!!] https://www.youtube.com/watch?v=EsOU0V2kpUI that seems to borrow on the theme of a computer ghost character.

  • Google Translate will hallucinate translations for complete nonsense, so this probably doesn't mean anything.

    • WWWJDIC/JMdict also claims it's a name "Junko", but it also isn't very reliable. If you want to know what a Japanese word means you should look it up in a JP-JP dictionary.

It looks as if these (at least 妛) are being used in various places on and offline. It’s eventually possible that they will become associated with one or more meanings and perhaps a pronunciation.

  • In East Asian cultures that use Han characters, people used to make up new characters when the need arises.

    These days, we scroll though the Unicode standard and find rarely used characters that were accidentally added and imbue them with new meaning. (yes, this is seriously a thing)

    • When the article said:

      "In the end only one character had neither a clear source nor any historical precedent: 彁."

      my instinct was that this character could be retconned to mean "character whose meaning has been lost", thus creating a self-referential paradox.

      Presumably someone would have to then separately come up with a pronunciation for it. Perhaps pronouncing it "duangu" would solve another problem:

      https://coconuts.co/hongkong/lifestyle/duang-jackie-chan-ins...

    • Oooh that sounds fascinating. Any examples of that that spring to mind? Is the pronunciation (or a reasonable representation thereof) already recorded in the Unicode standard or is that also a bit of free-jazz?

      4 replies →

    • One of the reasons I wish a compositional language had been standardized for Unihan instead of the code-point-for-every-character approach.

      1 reply →

This is a tangent, but I felt like sharing. In college, I purchased a used copy of the communist manifesto. Famously, the first line reads, "A spectre is haunting Europe, ...".

The previous owner had both highlighted and circled the word "spectre" and wrote "ghost?" in the margins. The rest of the text was similarly marked up.

Every time I hear the word "spectre" I see "ghost?" in my mind's eye.

I'm more worried about the inflation of emoji than a couple dozen unused ghost JIS characters.

  • Godwin's second law: any sufficiently long discussion about Unicode includes a discussion about emoji :)

  • Why? Unicode isn't running out of space any time soon.

  • If Slack/Discord/etc. custom emojis get used enough, do they get incorporated into Unicode? I've seen something like 40 variants of laughing emoji, and closer to 400 variants of Pepe the Frog, and I'm not even in any "alt right" or 4chan-adjacent chat rooms/guilds where I imagine there are even more. Not to mention the countless custom anime face ones.

Many (too many) online forms in Japan are unable to process foreign names, sometimes for reasons as brain dead as not allowing more than a few characters in length.

Even a perfect unicode standard wouldn't be able to mitigate the arrogance of a programmer.

Summary:

Some Japanese characters that aren't real got accidentally built into the unicode code table. This is NOT related to speculative execution attacks at all. It's just "whoa, these kanji don't mean anything, how the hell are they here?!"

Other than that caveat (not security related), this is a fascinating article, especially if you've studied Japanese as a (foreign) language.

Z̵̘̋̎̕ả̶͓͑l̶̜̈͒g̸̡̧̤̋͆õ̶̡͔̥̓ ̵̱͌̈́͝ẃ̴̫̤͘a̶͈̭̱͒̊i̴̭̪̾͑̕ṫ̵͈͙̻s̵̺͐̅̊ ̶̮̩͒̋

Is there anything similar for Latin characters?

The only circumstance I can imagine is where a Latin character has been erroneously encoded with an unused diacritic, for instance a T with a diaeresis.