Comment by gucci-on-fleek

1 day ago

> Not that hard to imagine, OCR existed back then?

How do you train OCR on a character that doesn't exist though? Even these days, OCR often confuses the 52 Latin characters, so I wouldn't expect as good of results with older technology and some 20k CJK characters.

(Talking about Japanese here)

I don't know if this is how OCR would work here, but they're made of components called radicals. It's kind of like how letters make words, but in a two dimensional layout, and one letter was wrong and a nonexistent word was made.

For example 彁 is made of parts of 引 and 歌 (the first that popped into my mind).

IIRC there's only like 200-something radicals, though some can be hard to spot due to how they overlap or nest inside each other.

  • I find 鸚哥(いんこ) to be the easiest way to type the right hand side of 彁.

It does exist? It's part of Unicode!

  • It's listed in the Unicode character database, but I doubt that it was in (m)any fonts 20 years ago, it would have been in zero printed books [0], and nobody knew the meaning of the character, so I'd argue that it's more an artifact of the Unicode compilation process than a "real" character.

    [0]: Aside from the "Overview of National Administrative Districts" book mentioned in the article

    • You don't need (m)any fonts, only need one font used by the book and you can get it the same way it was originally done - gluing some parts of other characters together. Or you could just draw it if the Unicode version is too dissimilar from the book version for a reliable OCR. (though if the character came from this single book, likely the Unicode reference is its replica?)