← Back to context

Comment by jiggawatts

7 days ago

Literally "just" what I outlined! That's the brilliant thing with LLM-based automation, you don't need a massive piece of complex software, just "ask" in English.

Roughly:

Get API keys for multiple vendors or just use OpenRouter (but availability of frontier models tends to be limited). Alternatively, Azure Foundry has everything except Google models, so just two subscriptions is enough.

Run the same prompt and same input image through each of your chosen models.

Then feed the smartest model the original image together with the collected output texts. Use a prompt along the lines of "Merge these attempts to OCR together into an corrected and improved combined version, taking special care to exactly preserve the original's typos, etc, etc..."

You can do this manually, it's just fiddly. It's not hard to automate, most of the "code" is English instructions!

The downside of this approach is the cost: even the "light" frontier models are a few cents per page, which is not so bad until you're doing this 5x or 10x times per page and suddenly scanning a notebook can set you back tens of dollars, more than buying a good novel at a book store.

I picked up on this technique back when GPT 4 was released. People noticed that it could translate ancient Akkadian, but only if you ran the prompt through 4x times and merged. I tried this with a few random samples I found online and the merged translations were generally better than the "official" ones, even thought the individual attempts were unreadable gibberish.

There are already scripts/tools floating around for this!

Look into OpenRouter Fusion, Consensus AI, Multi-Model Debate, etc... or just whip up something yourself.

Sounds kinda like exactly what I set up in the repo I linked in the original comment heh... Minus the openrouter calls.