Comment by vunderba
3 hours ago
Neat. What are your thoughts on the recently released Iris-3B, a pixel-space model which also bypasses the need for a traditional VAE?
3 hours ago
Neat. What are your thoughts on the recently released Iris-3B, a pixel-space model which also bypasses the need for a traditional VAE?
Thanks for links! I hadn’t seen this yet.
Very related; our models are quite literally cousins, as we’re both interesting on the work of JiT from last fall.
Architecture is roughly the same up to some minor differences (they use MM-DiT from FLUX, we use single stream from Z-Image).
The biggest delta is that we’re going with 32x32 token reduction and they go for 16x16. Our goal here was to cut attention windows 4x to speed up training inference. And have been publishing work (this included) on how to match models with less compression to than us.
In this blog specifically there are 3 things we stack on top of each other to get good results:
1) transformer blocks per modality before they enter shared single stream transformer blocks
2) very different noise schedules during training than those proposed in prior work
3) predicting the output at k different resolutions along the depth of the diffusion transformer to provide gradient signal earlier in the network to learn better (and faster)
The other work does a great job scaling up the JiT to 3B parameter; this is purely an experimental release highlighting how we make up the gap.