← Back to context

Comment by acgourley

4 hours ago

Very cool.

I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.

Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.

>Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context. Very true in my case on similar problems, my major issue was OCR relics. Reasonable mispelled words say by an uneducated person, are not that much of an issue. For the OP VOC work most letters were written by educated scribes and less of a problem. Anything before 1650 had very different calligraphy though.

Please just make sure to keep the ethical implications of any such work in mind.

I do not know what exactly it is you're building, but the shape also fits "weapon", and weapons do not really care about the good intentions of their author.

  • I hear you, I think on balance it's good which is why I'm working on it. It makes elite opinion legible and helps detect organized disinformation dark matter. It's not like the intelligence and advertising markets needs help from me about how to surveil downwards.