← Back to context

Comment by lmkg

4 years ago

One could argue they're facets of the same issue. Although in the spirit of the original list, they would probably get split into separate line items.

On further review, I think this is als similar to #12 & #13 on the list: "names are case-sensitive," and "names are not case-sensitive." To generalize that to include non-Western alphabets: display variations of the same character are significant, and display variations of the same character are not significant.

This of course goes back to the evergreen philosophical question "what even is a character, anyways?" Since we've found a case where two characters which are the same character are not the same character. Are they distinct characters or typographical variants? Yesn't: one would want them unified for searching, but distinct for printing.

But regardless of what they are, these characters/variants only show up in names. Names tend to retain archaic (or extinct) language variations longer than speech, which is the reason for rule #11, which is at least part of the problem.

I fully agree with this second, expanded take of yours. Some names are both represented and not represented by the Unicode simultaneously. This suggests there should be variant versions of characters, but that becomes an even thornier combinatorics (and sorting/collation, and lookalike characters) issue than what already exists.