hn.today

UniEvo-VL: Self-Distillation Training for Multimodal Model Self-Improvement

arxiv.org11 points2 comments
Screenshot of UniEvo-VL: Self-Distillation Training for Multimodal Model Self-Improvement

UniEvo-VL introduces an on-policy self-distillation training recipe that lets a single multimodal model improve itself by learning from its own critiques during test-time compute. The method treats the model as both teacher and student under different contexts: the student receives the original prompt while the teacher conditions on privileged self-critiques, and training minimizes per-state divergence between their denoising-diffusion distributions along the student’s sampling trajectories. This avoids reliance on an external, larger teacher and leverages internal reflection as privileged information, applying self-correction directly to the model’s generative process.

Empirical results build on Qwen-image-2512 and show measurable gains in image-generation evaluation: GenEval rises from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53. Experiments with stronger external critics such as GPT5.6-Luna indicate that models with better internal judging ability can reach a higher self-evolving ceiling. Gains are concentrated in image generation while sensitivity to additional reflection signals is retained, but improvements are task-dependent - mixed text-rendering results demonstrate non-uniform benefits. The approach frames a practical, compute-on-test-time route toward recursive multimodal self-improvement without external supervision.

Read on arxiv.org2 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.