Comment by hnfong
4 years ago
This might be interesting read to those unfamiliar with CJK, but character bloat(?) isn't remotely a recent thing. It's actually at least a couple hundred years old.
The Kangxi dictionary (1716), an authoritative dictionary of Chinese characters, contains definitions for 47035 characters, even though only a couple thousand are in common use. Quoting from Wikipedia: "The dictionary was the largest of the traditional dictionaries, containing 47,035 characters. Some 40% of them are graphic variants, however, while others are dead, archaic, or found only once. Fewer than a quarter of the characters it contains are now in common use."
All of these archaic (or even bogus in some cases) characters found in the dictionary are now part of the Unicode standard, of course :) The unihan database even has a field that shows the page number where the character appears in the Kangxi dictionary. If you're wondering why 65536 characters isn't enough for everyone, the junk in Kangxi dictionary is a significant contribution.
Unicode is a mistake that could only have happened in turn of the century America.
It is the distilled essence of the idea that you need to be inclusive of everyone along with a fundamental ignorance of what anyone who isn't American does.
The idea that Chinese characters are glyphs in the same sense of Latin characters can only have come from someone who has never written Chinese.
It is as stupid as demanding a glyph point for each possible integral, e.g. https://quicklatex.com/cache3/8c/ql_9739884527bd893429657272... and https://quicklatex.com/cache3/18/ql_c51509950f58a52253c696a4....
Solutions for English are not solutions for all languages. You can tell because the solutions that were natively invented by people who spoke those languages were _not_ unicode.
JIS, Big5, UHC, and GB all use a codepoint-to-character approach. You're right to point out that many aspects of "multilingual" support are written by people who do not know another language and so end up being hopelessly misguided but it's not really fair to say that Unicode invented this and thrust it upon the CJK world. Every pre-existing system of representing 漢字 had a codepoint table (which Unicode references in their description of each character).
Han Unification was in my view problematic but was driven by technical limitations (then again, if Simplified Chinese characters had also been unified I suspect there would've been more pushback to come up with a better solution, but ultimately Japanese was stuck with being the only one making a major compromise on that front).
I don't think a stroke based or combination system would've been better for many reasons: https://news.ycombinator.com/item?id=32102093. And if you don't trust Americans who at least tried to learn about the subject matter, how much do you trust any other programmer (who has no interest in other languages) to be able to handle a more complicated system for representing and rendering 漢字?
I don't know about Chinese since I'm barely literate in it.
But copying how every Chinese dictionary renders characters as trees of simpler characters seems like a much better approach than the arbitrary letter to number mapping of unicode. That it works well in English is only due to the fact it has no accents.
I can confidently say that native solutions for scripts with accents were _not_ unicode like but overstrike. My grandfather has the source code, in Romanian, of a 1960s computer payment system he worked on which had to deal with both Romanian and Hungarian names.
The combinatorial explosion of possible letters and accents made unicode like encodings an obvious non-starter. Historically names could pick up any accent (some times more than one) on any letter. When you have 26 base letters and 6 possible accents you'd need 26 + 26 * 6 (182) unique representations for single accented letters and 26 + 26 * 6 + 26 ^ 2 (1118) for double accented letters. That makes a language which is fundamentally alphabet based unusable on a keyboard. Something that Unicode is still sweeping under the rug.
3 replies →
> Every pre-existing system of representing 漢字 had a codepoint table (which Unicode references in their description of each character).
Even typewriters worked with a giant table: https://en.wikipedia.org/wiki/Chinese_typewriter
Han Unification was invented by Chinese people (in Hong Kong iirc), not by Americans. And even before unification, the national standards also had one code point per kanji/hanzi rather than building them up out of radicals, as opposed to the way eg flag emoji are done.
JIS did this in 1978 for instance.
It does appear that computerizing CJK languages has made them very different from handwriting them; native Chinese speakers now constantly forget how to write hanzi. But they did this to themselves.
> It does appear that computerizing CJK languages has made them very different from handwriting them; native Chinese speakers now constantly forget how to write hanzi.
This is separate to the question of encoding -- phonetic-based input systems (IMEs) are a far more likely cause (there are less-widely-used shape-based IMEs which still spit out a Unicode codepoint). The same is happening to Japanese natives, though it should be noted that it's not the case that they cannot write 漢字 normally, they just might forget how to write a relatively rare one (just like how you might forget how to spell a word in English because of a dependence on autocorrect and spellcheck).
I think 'character bloat' is simply inherent to the writing system when characters are written by hand (now that perhaps most written communication is digital people can't use characters that are not already supported)
Anyone can invent characters whenever they want, and it's only a question of them sticking or not.
I think this is also one of the reasons for the Chinese tendency to push for unification and uniformity.
When it’s character based instead of alphabet based, I think it’s the equivalent of coming up with a new word in English, which is basically what you’re describing.
Sometimes it’s mashing two previously unrelated ‘words’ together (aka the tons of compound characters in Chinese), other times it’s coming up with something completely new.
Same rules apply though, if it doesn’t add value worth the trouble (or get mandated by the powers that be), it’ll eventually just die out or be a curiosity.
Also, to keep it tech related:
RISC = English CISC/VLIW = Chinese?
IIUC Old Chinese was a much more “isolating” language, in that words were typically single characters - meaning that to make new words, you typically needed to make new characters. As it evolved through the ages, “compound” words composed of multiple characters became more common. These days, new words are almost always combinations of multiple characters (often 2, occasionally 3-4).
7 replies →
> Sometimes it’s mashing two previously unrelated ‘words’ together (aka the tons of compound characters in Chinese), other times it’s coming up with something completely new.
That's not how it works. Most Chinese characters stem from a character C having a pronunciation A referring to a meaning M being used to note another word of meaning M' with same pronunciation A (sometimes slightly different A'). This of course doesn't scale really well, hence the existence of determiners in logographic scripts, which are words used without their pronunciations placed before or after another to give a semantic clue. The innovation of Chinese (which I think is why it's still an efficient script today) was to incorporate the determiner in the character itself to give birth to a character C' where a part refer to the pronunciation and another acts as the determiner, instead of padding the main text with (a lot of) determiners.
I'm not sure I understand. Most European languages go through cycles where letters are added when languages are mixed together followed by periods of redundant letters disappearing. Old English had something like 39 letters. 'th' used to have its own letter: thorn.
I think character proliferation in CJK languages are a result of each word having its own character. The proliferation isn't fundamentally a proliferation of characters, it's a proliferation of words, which happens all the time in all languages. But only in certain languages does this proliferation of words result in additional characters being added to the language.
> I'm not sure I understand. Most European languages go through cycles where letters are added when languages are mixed together followed by periods of redundant letters disappearing. Old English had something like 39 letters. 'th' used to have its own letter: thorn.
There is a fundamental difference between pictographic languages where glyphs have intrinsic meaning, and alphabetic languages where letters reflect sounds.
Old English was much better spelled than current English because it didn't have a spelling. People wrote what they heard. The current mess is because we have 5 centuries of bad standards that can render ghoti as fish. I think you'll agree that the digraph ti, as in nation, is just as nonsensical as sh for the same sound and we'd be much better served by a single glyph for both.
We in fact have that already: https://en.wikipedia.org/wiki/International_Phonetic_Alphabe... English uses somewhere around 45 of those sounds depending on accent. Th for example renders two distinct sounds: ð and θ. þ is not any better than th, apart from brevity, because it also rendered to ð or θ when spoken depending on context.
Chinese is of course as much a pictographic language as English is an alphabetic one. A substantial number of glyphs come from combinations of simpler glyphs which have the same sound as the word you're trying to write.
>Fewer than a quarter of the characters it contains are now in common use
12K characters in common use is equally impressing for me as a non-Asian.
It's actually way fewer than that IRL. Japan's official list of commonly used Kanji only has 2136 characters. Taiwan's list has 4808, and the PRC's list has 3500 "frequent" characters with another 3000 supplementary "common" ones. Digitization has made it even easier to use these characters without recognizing the actual form or how to write them.
The 常用漢字 (Japanese Common Use Kanji) list does not include many kanji that native speakers can read and newspapers don't always follow the rule that they only should use characters from the list. In addition, you need to include the 人名用漢字 (Personal Name Use Kanji) in the list because basically all of those characters are also used in fairly common words.
Native speakers can probably recognise at least 3-4k kanji if not more but can probably only write around 2k from memory, depending on how well-read they are.
嘘 (lie) is the best example of an incredibly common word whose kanji form (which is used fairly often) is not in any official government list.
If you look at a frequency list of Chinese characters,[0] the top 4800 characters make up about 99.9% of modern texts.
That means that if you know 4800 characters, and you read a text that is 1000 characters (equivalent to around 700 words) long, there's likely one character you won't recognize.
The funny thing is, if you recognize only the top six characters, you already know 10% of the characters in a typical text. The distribution is very top-heavy, but with a long tail that you do have to learn to become literate.
0. https://lingua.mtsu.edu/chinese-computing/statistics/char/li...
2 replies →
More like 12k characters currently in use at all. Common use characters are a much smaller set than that. (3k or so?)
Does Unicode really need to store Chinese words? Is it impossible to deconstruct the glyphs into strokes, each stroke effectively being a character?
The problem with that would be that every software must know the intricate rules about combining glyphs, and if they guess wrong, users get garbage characters.
Considering that the majority of code is written by people who don't know Chinese characters, it would result in never-ending issues, pretty much everywhere.
Korean actually has a two-way system in Unicode. Every conceivable character (= syllable) possible in modern Korean has its own codepoint, which allows most software to display them correctly: from their point of view, it's just another CJK character.
On the other hand, there is a Unicode area containing Korean sub-blocks ("jamo") that were used historically. In theory, you can combine them and get some pretty funky archaic syllables. Almost no software renders them right.
They can't even get much simpler things right. Qt incorrectly combines accents with the character to the right instead of the left and has been refusing to fix this bug for years.
In addition to the problems mentions by yongjik, even with the current system, very little software is even aware that the same codepoint should be rendered differently in different languages (返す needs to be rendered differently in every CJK locale) which often results in websites and programs using Chinese fonts for Japanese text (even if you've configured your language as Japanese). Having stroke breakdowns would not make this situation better because there are multiple ways to render the same stroke description and there aren't really systematic rules for how to correctly represent the Japanese (or Taiwanese or Korean) version of a character -- it's generally for historical reasons. If you were to try to actually represent the characters faithfully (in an attempt to avoid making every country unhappy with the way you've butchered their language), many characters would become unusable for text searching because the same "character" (from the perspective of a CJK native) would have a completely different representation in a way that a computer could not be able to identify as being the same (even a character as simple as 言う would have this issue).
I dread to think what an enormous mess would result if every character was represented as a build-it-yourself instruction manual rather than allowing font authors to correctly represent the characters. This is also ignoring that (depending on the font style), the apparent strokes for a character can change between fonts in the same language (this is because the computer font stroke style and the written font stroke style can be different) -- by putting stroke decisions in the encoding you're introducing a layering violation since fonts should be deciding how characters are styled, not encoding format committees.
Also nobody in China, Japan, nor Korea would switch to an encoding system so incredibly inefficient that more strokes results in more bytes being necessary to store the character (they already compromised with having 3-byte UTF-8 characters when JIS, GB, and Big5 all only required 2 -- and Japan was basically forced to compromise on Han Unification). This would've resulted in the failure of Unicode's mission to be the One True Encoding Format.
In the early days of computers some character systems were stroke-based because that used less memory than a 32x32 bit map. A kilobit of ROM (one character) could cost $10.
Currently stroke-based systems are used for calligraphic effect. You could generate new font types, e.g. bold., but controlling the shape of strokes.
Stroke systems are important for teaching character writing because the drawing order is rigorously prescribed. Once you learn the first couple hundred, you can pretty much guess future characters. Wrong order characters often look bad and suggest a non-Chinese speaker mis-copied them. (e.g. some tattoos)
Unicode has support for this, in the Ideographic Description Characters block (https://en.m.wikipedia.org/wiki/Ideographic_Description_Char...). However, it’s purely descriptive, and not designed for rendering.
There are somewhat more sophisticated systems which define both the rendering and stroke decomposition of characters (e.g. CDL: http://guide.wenlininstitute.org/wenlin4.3/Character_Descrip...). The general workaround for characters that aren’t on Unicode would be to use one of these stroke description systems to create the character, then render it to an image and insert it.
Many attempted, but nobody have suceed. The most famous one is `Chu, B.F.: 漢字基因朱邦復漢字基因工程 (Genetic engineering of Chinese characters) (2003), http://cbflabs.com/down/show.php?id=26 `