norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers55 minutes ago

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

The paper introduces Iris-3B, a 3B-parameter pixel-space text-to-image transformer pretrained through a 256to512to1024 curriculum. It compares pixel-space models with latent models in tasks like depth estimation and image restoration, finding no significant improvement in performance. The study also converts a latent model to pixel space and evaluates both approaches.

Open original
SIGNAL FROM THE SOURCE
6
source votes
Tracking sinceOctober 8, 20266 source votes
Momentum+1.64/hover 3.05 h
Discussion—Read comments ↗
PublishedOctober 7, 2026Hanqiu Li Cai, Chema Garabito
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

6 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research provides insights into the effectiveness of pixel-space models compared to latent models in specific tasks, which could inform future work on image generation techniques.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic