← Back to context

Comment by knollimar

11 hours ago

Oof 800 by 800 kills a lot of use cases

Might still be fine. The most recent crop of vLLMs proactively use whichever programs are available on the system (e.g. ImageMagick or PIL) to "zoom in" by cropping subimages if they can't quite make out the details.

  • Downsizing a higher res image to lower res means the zoom will be blurry.

    • They process the original image file with Python on the local device. (And I've seen the web chats do this with their "computer use" features too.)

      The really wild one is even blind models will do this and they'll try to run stats on the pixels to figure out what it looks like... the even wilder thing is that it kind of works!

      2 replies →

    • The order is:

          LLM issues tool call to read high res image ->
          harness sends high res image to server ->
          server downsizes it to 800x800 (blurry) ->
          LLM issues bash command (e.g. `convert`) to crop a small subimage (e.g. 600x600) from the high res image ->
          LLM issues tool call to read subimage ->
          harness sends subimage to server ->
          server does not resize the subimage because it is small already, so it is not blurry when finally ingested by the LLM

      2 replies →

For most use cases you can fix that in the harness. Just give the model a tool to request a crop of specific coordinates of any image it has in its context. Call the tool "zoom" and it should be intuitive for the model

Maybe there are some use cases where you need high detail everywhere at once, but for OCR of small text and the like a zoom ability should be sufficient

  • For really dumb models I've also had success automatically cropping it into a grid of N images with the max size, then processing each cell individually, then once all been processed, do one final call with resized image + all other context previously generated per cell. Basically a workaround to the image dimension restrictions without loosing fidelity. Works well with even dumb 7B models.

    Can't remember if I stole this idea from some existing public harness though, can't remember. If someone knows of public harnesses that do this already, please share them :)

I don't know about a lot. Probably more like a few. I take a lot of screenshots for various reasons, and over 800 seems like I could have done a better job framing and cropping.

It might also be due to its experimental status. Wouldn't surprise me if the GA version allows for larger input. Either that or the eventual pro version.