Pointing AI at archives found a forgotten meteorite, lost rhinos, and more

10 hours ago (jessewaites.com)

If I sat down to read just the Dutch East India Company pages myself, at two minutes a page, eight hours a day, five days a week, it would take me about 70 years. And that’s before the newspapers. My homebrew AI lab got through the entire archive in a single twelve-hour overnight run.

Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...

  • Surely more than if they read nothing at all because most of it would be tedious drudgery of minimal value.

  • I'm sure that if he didn't do what he did, he would lear about Dutch East India Company so so very much.

    Remember that even junk food is more food than junk and you can survive on it for years.

  • Probably a lot. I’ve never learned about so many disparate subjects as I have over the last two years with LLMs.

    Why is it always these supremely weak arguments and rationalizations against LLMs that come from people that have been intelligent, at least based on their comment histories, for so many years. It’s radicalizing me. I want a data center everywhere and I want tokens to be so cheap they’re like electricity or water.

    • At school i was made to read all of shakespeare. It wasnt about passing tests. Any LLM can read shakespeare and pass a test. I was made to read shakespeare so that i would appreciate language in the hope that i would strive to improve my own. That still counts.

      I am a soldier and language is vital in every day of my job. Tone, word choice, cadence ... all of it conveys meaning in a way that an LLM simply cannot. Soldiers will not follow an AI up a hill, nor will they follow those they know use AI to pretend they understand a subject.

      Want to sound smart? Watch blackadder. Want to win no-win aguements through wit? Watch Archer. Want to inspire your troops to do somthing unpleasant? read and watch shakespeare. An LLM can teach you nothing that really counts.

      1 reply →

  • You know you don't have to comment on everything you see on the internet, right? You're allowed to just keep scrolling when something isn't for you. I wonder how much joy you have in your life, I suspect very little, if anything.

    • I didn't mean to upset you, piratebroadcast. Like so many of us, you scratched a technological itch, and that is fine. However, now that AI tools have supercharged us, these scratches are easier to scratch, and their outcomes, made public, flood the aether.

      1 reply →

    • You're missing the actual point sorokod has.

      Specifically that these rabbit holes are useful to bring people to all sorts of new discoveries and skills.

      It is important to point out that that's a real risk with such AI use. Of course, it is also true that it would likely not have happened at all otherwise. Both things can be true at the same time.

      __

      Oh I just realized that you're the actual author and this is not the only super-thin-skinned comment.

      Man. Why do be like this.

      3 replies →

>To make this kind of research accessible, I’m open-sourcing the workflow I created for this investigation as a small toolkit, Antiquity, enabling anyone with a question and a coding agent to conduct similar historical archival investigations.

https://github.com/jessewaites/antiquity

  • There are also VOC archives at Cape Town, also in Kew (search for the letters of Loot) which were literally looted by privateers. All these are written in High Dutch some in German. How reliable are the translations?

IMHO: The rotating rhino, meteor impact, and animated flowchart is totally unnecessary cruft that makes it look almost satirical. If this keeps up, in time, this "AAA effects" stuff is going to look like the 90s "under construction" banner gifs.

  • The effects are comically bad. I see the inspiration in scrolling effects that the New York Times put together, but the NYT was never dumb enough to obscure the copy text. Form follows function, and the function of a web page is to be read, not to obscure what is to be read with some stupid effect that's supposed to remind one (I suppose) of a volcano's cloud obscuring one's vision. At least "under construction" banners didn't obtrude upon the copy text.

  • I mostly agree, I still read and enjoyed it but part of me kept snagging on the fx and wishing it was way way turned down.

That was a fascinating read, I really enjoyed it. Literally like exploring lost knowledge. Great work and a great write-up.

I also liked the aesthetics of it and the little effects (meteorite and volcano, but please fix the rhino and the text flowing around it while it rotates).

I wonder what else could be found in such archives. Some ideas: - Locations or routes of sunken ships and their missing cargo?

- Some pirate stories, maybe about a now-forgotten but once-legendary pirate captain?

- Unusual weather events, like snow in the summer?

(edit: formatting)

Very cool.

I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.

Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.

  • >Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context. Very true in my case on similar problems, my major issue was OCR relics. Reasonable mispelled words say by an uneducated person, are not that much of an issue. For the OP VOC work most letters were written by educated scribes and less of a problem. Anything before 1650 had very different calligraphy though.

  • Please just make sure to keep the ethical implications of any such work in mind.

    I do not know what exactly it is you're building, but the shape also fits "weapon", and weapons do not really care about the good intentions of their author.

    • I hear you, I think on balance it's good which is why I'm working on it. It makes elite opinion legible and helps detect organized disinformation dark matter. It's not like the intelligence and advertising markets needs help from me about how to surveil downwards.

Okay, for those of you that do not enjoy fun, I just added a button at the top of the page to turn off the special effects.

  • I honestly primarily enjoy my browser performance not tanking on this $5000 workstation and actually bailed out while scrolling because it wasn't really usable.

    I bet it works great in chrome tho

I love this, one major friction I’ve seen to human coordination and advancement has been the journals in different languages

Many people don’t notice, but even Wikipedia has no normalization between articles in different languages. The language button there acts like its showing you a translated version of the article but its actually a completely different Encyclopedia and community of editors with no cross reference to the other language’s article and references at all. Articles that are stubs on the English page may be massive fully fleshed out articles in another language, and nothing native to the site or anything I’ve seen will tell you that there is more information in one variant

LLM’s can find the word associations and compare them in all languages, even if it itself doesn't innately know language

and there would be so much low hanging fruit here like this engineer found