Comment by jamienk

8 days ago

My dad died and had many many notebooks of his journals with very hard-to-read handwriting. Is it worth the effort to scan all of these so I can feed them in and go to work. Seems like so much minutia is out there, ready to be meta-understood.

I set up an OCR flow using local models on all my many tens of journals stretching back the last 30 years.

I would say it's about 80% accurate, which means it's missing enough key words to make a lot of it uselessly unintelligible. I can easily compare the images against text I turn up in a grep which is nice if I'm looking for something.

Allegedly Claude set up a system for retraining for my handwriting, but it would require me to manually revise several hundred pages by hand so I don't think I'll ever do it.

https://github.com/508-dev/journal-ocr

  • Accurate text OCR from bad handwriting is still very much a "bleeding edge" frontier capability that isn't practical with local models.

    GPT 6.1 and Gemini Flash 3.8 both do pretty well, their OCR of your sample image is only "wrong" in the sense that the original has typos and they corrected some inadvertently and/or filled in gaps where you had "unintelligible" in the canonical text.

    If you have the budget and want the best possible results, you need to run each image through multiple models and then combine the outputs into a final "merge these" prompt. Better scanning helps too, your sample image is rotated and you used a phone in low light. Try a DSLR or a flatbed scanner and process only one page at a time instead of two at once.

    • > you need to run each image through multiple models and then combine the outputs into a final "merge these" prompt.

      I haven't tested this recently but my possibly dated experience is frontier LLMs can't figure out which model is correct or incorrect if there's disagreement on vision recognition.

      Have you found otherwise?

      (Edit: I see you gave an anecdote about merging terrible results. My experience is with merging overall accurate results).

      1 reply →

    • 'run each image through multiple models and then combine the outputs into a final "merge these" prompt' << How to do this? This would be an amazing workflow to get documented. This could be the start of a full-service company "send us a bunch of notebooks, get back HIGH QUALITY text version"

      3 replies →

I think scanning is worth it. There's a lot of cheap services out there that do it but it's also not too hard to find relatively cheap scanners that will process stacks of pages if you are willing to destroy the binding.

I scanned quite a few documents many years ago and it's been fun trying new tools every couple of years to see how good they are getting. I would say it's still not 100% there, but if you have them scanned you can basically just just keep trying and compare the results.

I think the only disappointment in recent years has been that storage has gotten more expensive instead of cheaper. I was waiting for SSDs to get cheap enough to justify moving all my documents to a fast flash array to process and search through them faster, but that doesn't seem like it will happen anytime soon.

  • >I was waiting for SSDs to get cheap enough to justify moving all my documents to a fast flash array to process and search through them faster, but that doesn't seem like it will happen anytime soon.

    How much data are you working with? lol

    Seems like you're doing something professional-grade if this is the case. I imagine most people, like the OP, can basically use whatever machine they have laying around and never be concerned with storage size/speed.

    I'm also curious if you've tapped into cloud computing. Not that using the cloud is cheap, but I suppose if I was in a position where I'm concerned with the cost of SSDs for a processing task, then I'd be exploring all of my options and I'd be surprised if the cloud wouldn't be an "easy" solution.