The falsehood here is thinking that if you can encode the name into the right code points, and you have a font that can print them, the result will be acceptable to the people whose name it is.
They had that, but needed a font that used a different number of strokes for the characters because of the superstition.
More generally, the notion that human culture, systems, and behaviors can be mapped, losslessly and without causing harm, to something a computer understands.
I think these language examples are so good, as examples, because all aspects of them are clear and easy to follow. I think computerization of business and society and the systems that make them work, causes immense amounts of this kind of friction and pain all the time, in ways that are much harder to understand, explain, or catalog (which is precisely why it's such a big problem, though as far as I know it's received little attention)
[EDIT] To distill it, I think that trying to make a computer a "source of truth" rather than a tool, tends to do substantial violence to the "truth".
One could argue they're facets of the same issue. Although in the spirit of the original list, they would probably get split into separate line items.
On further review, I think this is als similar to #12 & #13 on the list: "names are case-sensitive," and "names are not case-sensitive." To generalize that to include non-Western alphabets: display variations of the same character are significant, and display variations of the same character are not significant.
This of course goes back to the evergreen philosophical question "what even is a character, anyways?" Since we've found a case where two characters which are the same character are not the same character. Are they distinct characters or typographical variants? Yesn't: one would want them unified for searching, but distinct for printing.
But regardless of what they are, these characters/variants only show up in names. Names tend to retain archaic (or extinct) language variations longer than speech, which is the reason for rule #11, which is at least part of the problem.
It's already there, #11 "People’s names are all mapped in Unicode code points."
The falsehood here is thinking that if you can encode the name into the right code points, and you have a font that can print them, the result will be acceptable to the people whose name it is.
They had that, but needed a font that used a different number of strokes for the characters because of the superstition.
More generally, the notion that human culture, systems, and behaviors can be mapped, losslessly and without causing harm, to something a computer understands.
I think these language examples are so good, as examples, because all aspects of them are clear and easy to follow. I think computerization of business and society and the systems that make them work, causes immense amounts of this kind of friction and pain all the time, in ways that are much harder to understand, explain, or catalog (which is precisely why it's such a big problem, though as far as I know it's received little attention)
[EDIT] To distill it, I think that trying to make a computer a "source of truth" rather than a tool, tends to do substantial violence to the "truth".
5 replies →
One could argue they're facets of the same issue. Although in the spirit of the original list, they would probably get split into separate line items.
On further review, I think this is als similar to #12 & #13 on the list: "names are case-sensitive," and "names are not case-sensitive." To generalize that to include non-Western alphabets: display variations of the same character are significant, and display variations of the same character are not significant.
This of course goes back to the evergreen philosophical question "what even is a character, anyways?" Since we've found a case where two characters which are the same character are not the same character. Are they distinct characters or typographical variants? Yesn't: one would want them unified for searching, but distinct for printing.
But regardless of what they are, these characters/variants only show up in names. Names tend to retain archaic (or extinct) language variations longer than speech, which is the reason for rule #11, which is at least part of the problem.
1 reply →