Most of today’s AI image generators get better the same way: they get bigger. More parameters means more memory, more energy and more expensive hardware, so quality stays tied to size. In our new paper, LiFT: Loop Flow Transformers, we try something different. Instead of stacking more and more unique layers, LiFT reuses the same block of layers several times in a row, a bit like an artist going back over a sketch to refine it. These generators start from random noise and turn it into an image in many small steps. At each step, LiFT can now “think” for a few extra rounds before deciding which way to go. Our motto: loop in depth, flow in time.

The key idea is in how the model is trained. Earlier looped models asked every round to produce the final answer straight away, and they tended to fall apart if you let them loop more often than they had practised. LiFT instead gives each round its own small task: move a set fraction of the way from the model’s first rough guess toward the correct answer. Because that “fraction of the way” is a continuous dial rather than a fixed count, a trained model can afterwards be run with far more loops than it ever saw during training, and it keeps getting better. One of our models, trained with just two loops, improved its image quality score (FID, where lower is better) from 17.1 to 9.3 when allowed sixteen loops, with no retraining at all.

The payoff is a small model that outperforms a big one. On the standard ImageNet benchmark, a mid-sized LiFT model produced better images than a conventional model more than twice its size, with roughly 60% fewer parameters, 32% less compute to train and 52% less compute to generate each image. That makes the approach attractive wherever memory is scarcer than processing time, such as drones, robots and other devices at the edge, and you can turn the number of loops up or down depending on how much time you have. There are honest caveats: at the smallest model size the conventional approach still wins, and so far we have tested only on 256×256 images of the ImageNet classes. Next we want to find out whether “think longer instead of growing bigger” also works for video and for models that simulate how the world moves.

Demo. Video.