Comment by harsh_patel14

16 hours ago

This is handy — I've hit this exact issue prepping documents for LLM context. How's the accuracy on lower quality scans?

Haven't tested rough scans honestly — my own use is rendered text, the easy case, where it's 93-95% confidence. I've been feeding it Kindle trading books into my finance app to compare strategies against my codebase, and an LLM is forgiving of the odd mangled word. Each page shows its confidence and a thumbnail of what was captured, so bad pages are obvious rather than silently wrong. Let me know how it does on low-quality scans if you try it.