← Back to context

Comment by user43928

4 days ago

I hope the open-weight versions will be SOTA.

> Over the next few weeks and months, we will make the following capabilities available

> Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)

> We will also release more technical details on the underlying approach.

Open-weight promise seems nice, but 1) I've seen people commenting that these hopes amounted to nothing for some previous releases (no idea which promises were made though); and 2) if the backbone is released as Dev, what will be missing? I can't easily tell from the post.

  • dev variants are usually cfg distilled which means that directly finetuning isn’t as effective. In the past, for the flux2 klein models,they released base versions that are not distilled. So it will probably be a while before you can fully take advantage of the open weights.

  • I run some of the 2.3 models locally so I'm not sure where you saw that they didn't follow their words?

    • It is not that they don't release open weights, but some users report that they are significantly inferior to the closed versions.

  • Could be handicapped version with more limited functionalities. Noentheless, open-weight is open-weight, I just hope the quality is good enough to be sota.

  • Flux2.dev and even klein 9b are extremely close to sota. People who are saying otherwise probably haven't used them very much.

    • I run a fairly high-traffic site for generative image models focusing on complex prompt adherence. Flux.2 doesn’t score anywhere near SOTA proprietary models.

      Klein 9b is decent for image-to-image, but when used for pure generative purposes brings back SDXL levels of body horror (have some Gattica pianists).

      If you can get past the annoying JSON structuring, Ideogram 4 is probably the best option in the open‑weights world right now for text2image purposes, and even scored higher than the original Nano-Banana.

      For ref: Flux.2 scored 5, Ideogram4 scored 8, and gpt-image-2 scored 12 out of 15.

      Comparison of Flux.2 [dev], Ideogram4, NB Pro, and gpt-image-2.

      https://genai-showdown.specr.net/?models=nbp,f2d,g2,id4

      2 replies →

  • Most of the disappointment around the previous open‑weight BFL release (at least for Flux.2 [dev]) came down to two main issues:

    • It initially required significantly more VRAM and was much slower than alternatives released around the same time (like Z‑Image Turbo).

    • The license felt overly restrictive.