Comment by mytailorisrich
4 years ago
I think 'character bloat' is simply inherent to the writing system when characters are written by hand (now that perhaps most written communication is digital people can't use characters that are not already supported)
Anyone can invent characters whenever they want, and it's only a question of them sticking or not.
I think this is also one of the reasons for the Chinese tendency to push for unification and uniformity.
When it’s character based instead of alphabet based, I think it’s the equivalent of coming up with a new word in English, which is basically what you’re describing.
Sometimes it’s mashing two previously unrelated ‘words’ together (aka the tons of compound characters in Chinese), other times it’s coming up with something completely new.
Same rules apply though, if it doesn’t add value worth the trouble (or get mandated by the powers that be), it’ll eventually just die out or be a curiosity.
Also, to keep it tech related:
RISC = English CISC/VLIW = Chinese?
IIUC Old Chinese was a much more “isolating” language, in that words were typically single characters - meaning that to make new words, you typically needed to make new characters. As it evolved through the ages, “compound” words composed of multiple characters became more common. These days, new words are almost always combinations of multiple characters (often 2, occasionally 3-4).
> These days, new words are almost always combinations of multiple characters (often 2, occasionally 3-4).
Yep! For example, the most common Chinese term for "Internet" is 因特网. This is composed of three characters:
互: "mutual"
联: "join", "coupled", "allied"
网: "net" -- carrying both the meaning of a woven net and a computer network
3 replies →
Any idea if it was due to things like the Confucian Official’s exam system (and corresponding increase in prioritization of education)?
More complex characters require more education to understand is my guess. Some of the traditional ones are….. obscure, and crazy complex.
2 replies →
> Sometimes it’s mashing two previously unrelated ‘words’ together (aka the tons of compound characters in Chinese), other times it’s coming up with something completely new.
That's not how it works. Most Chinese characters stem from a character C having a pronunciation A referring to a meaning M being used to note another word of meaning M' with same pronunciation A (sometimes slightly different A'). This of course doesn't scale really well, hence the existence of determiners in logographic scripts, which are words used without their pronunciations placed before or after another to give a semantic clue. The innovation of Chinese (which I think is why it's still an efficient script today) was to incorporate the determiner in the character itself to give birth to a character C' where a part refer to the pronunciation and another acts as the determiner, instead of padding the main text with (a lot of) determiners.
I'm not sure I understand. Most European languages go through cycles where letters are added when languages are mixed together followed by periods of redundant letters disappearing. Old English had something like 39 letters. 'th' used to have its own letter: thorn.
I think character proliferation in CJK languages are a result of each word having its own character. The proliferation isn't fundamentally a proliferation of characters, it's a proliferation of words, which happens all the time in all languages. But only in certain languages does this proliferation of words result in additional characters being added to the language.
> I'm not sure I understand. Most European languages go through cycles where letters are added when languages are mixed together followed by periods of redundant letters disappearing. Old English had something like 39 letters. 'th' used to have its own letter: thorn.
There is a fundamental difference between pictographic languages where glyphs have intrinsic meaning, and alphabetic languages where letters reflect sounds.
Old English was much better spelled than current English because it didn't have a spelling. People wrote what they heard. The current mess is because we have 5 centuries of bad standards that can render ghoti as fish. I think you'll agree that the digraph ti, as in nation, is just as nonsensical as sh for the same sound and we'd be much better served by a single glyph for both.
We in fact have that already: https://en.wikipedia.org/wiki/International_Phonetic_Alphabe... English uses somewhere around 45 of those sounds depending on accent. Th for example renders two distinct sounds: ð and θ. þ is not any better than th, apart from brevity, because it also rendered to ð or θ when spoken depending on context.
Chinese is of course as much a pictographic language as English is an alphabetic one. A substantial number of glyphs come from combinations of simpler glyphs which have the same sound as the word you're trying to write.