Pyramid-JiT (P-JiT) is a decoder-only, pixel-space text-to-image architecture designed to accelerate learning and reduce attention costs by predicting images at multiple resolutions along a diffusion transformer (DiT) trunk instead of using a VAE+encoder. The approach replaces a separate encoder with K small MLP output heads that regress intermediate low- and mid-resolution drafts (an inverted pyramid: 128→256→512), leveraging x-prediction so intermediate blocks produce useful low-res supervision without running expensive VAEs. That design concentrates parameters in a scalable decoder-like trunk, incorporates prior JiT improvements (wider patchification bottleneck, RMS-Norm, tanh gates, truncated AdaLN), and preserves dual-stream refinement innovations from earlier work while avoiding encoder-decoder parameter tradeoffs.
Empirically, P-JiT reduces loss and speeds convergence: the 3-readout pyramid cutting final v-space MSE by ~10% versus a single readout, yields visibly more natural details (faces, fur), and trains faster than JiT-DDT despite using higher-resolution patchification. It matches a previous Linum v2 FD-DINOv2 quality point in 11.3× fewer samples and 4.3× fewer GPU-hours at four times the pixels, with intermediate runs showing 2.18B active DiT parameters, 256 image tokens and competitive sample quality after ~75M samples. Code and weights are released under Apache 2.0 for research use.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.