hn.today

Training Text-to-Image Models Without a VAE

linum.ai3 points6 comments
Screenshot of Training Text-to-Image Models Without a VAE

Pyramid-JiT (P-JiT) is a decoder-only, pixel-space text-to-image architecture designed to accelerate learning and reduce attention costs by predicting images at multiple resolutions along a diffusion transformer (DiT) trunk instead of using a VAE+encoder. The approach replaces a separate encoder with K small MLP output heads that regress intermediate low- and mid-resolution drafts (an inverted pyramid: 128→256→512), leveraging x-prediction so intermediate blocks produce useful low-res supervision without running expensive VAEs. That design concentrates parameters in a scalable decoder-like trunk, incorporates prior JiT improvements (wider patchification bottleneck, RMS-Norm, tanh gates, truncated AdaLN), and preserves dual-stream refinement innovations from earlier work while avoiding encoder-decoder parameter tradeoffs.

Empirically, P-JiT reduces loss and speeds convergence: the 3-readout pyramid cutting final v-space MSE by ~10% versus a single readout, yields visibly more natural details (faces, fur), and trains faster than JiT-DDT despite using higher-resolution patchification. It matches a previous Linum v2 FD-DINOv2 quality point in 11.3× fewer samples and 4.3× fewer GPU-hours at four times the pixels, with intermediate runs showing 2.18B active DiT parameters, 256 image tokens and competitive sample quality after ~75M samples. Code and weights are released under Apache 2.0 for research use.

Read on linum.ai6 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.