Comment by hasibzunair

1 year ago

What kind of VLMs are being used in OmniAI?

I fine-tuned a Llama 3.2 Vision on a small dataset I created for extracting text without heavy cropping. Results are simply amazing in comparison with OCR-based approaches. It can be tried here: https://news.ycombinator.com/item?id=43192417