Comment by staticman2
8 days ago
> you need to run each image through multiple models and then combine the outputs into a final "merge these" prompt.
I haven't tested this recently but my possibly dated experience is frontier LLMs can't figure out which model is correct or incorrect if there's disagreement on vision recognition.
Have you found otherwise?
(Edit: I see you gave an anecdote about merging terrible results. My experience is with merging overall accurate results).
They can't, which is why I take the average of all results, and then have a final sanity check that considers the context, so if the text contains a bunch of references to "The Disposessed," and two models record "Shevek thinks Sabul is a propertarian," and three record "Shrek thinks Sabul is a propertarian," the final sanity check records Shevek.
I don't think this is viable for critical record OCR. I think the only way to do that is one pass with a frontier model and then a mechanical turk manual review passthrough with good compensation that allows for a slow and methodical approach. Plus of course much better scanning than a phone camera.
I basically kept rabbit holing this problem and finally settled on "80% and done is better than sitting on this problem for 4 years waiting to have time and equipment for a 99.99% solution." Crank the pictures between Claude code sessions, run the local LLMs when I'm asleep, done, now I can free text search years of journals plus I have photo backups now finally of them.