Self-Criticism as a Training Method: UniEvo-VL Lifts GenEval Scores Without an External Teacher
A new arXiv paper describes a training method in which a multimodal model improves its own image generation by criticizing itself — with author-reported gains on GenEval and GenEval2 Soft-TIFA. But the improvement is not even across tasks.
A research group led by Fang Wu, with 18 co-authors including Jure Leskovec and Yejin Choi, submitted the paper "UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement" to arXiv on September 30, 2026 (2609.38721). The paper describes a self-distillation method in which a single multimodal model acts as both teacher and student — using its own self-criticisms as "privileged information" to get better at generating images.
What's new
In conventional distillation, a smaller "student model" learns from a separate, larger "teacher model." UniEvo-VL breaks with that setup. "Rather than relying on a separate, often larger, teacher, we exploit their self-criticisms as privileged information and ask a single multimodal model to act as both teacher and student with different contexts," the authors write in the abstract.
The method is built on top of Qwen-image-2512, an open-source image synthesis model. The work is a concrete, measurable contribution to the research line on recursive self-improvement — but it concerns improving image generation on specific benchmarks, not a demonstrated general self-improvement capability.
How the method works mechanically
Rather than letting an external model judge and reward, UniEvo-VL lets the model itself generate criticism of its own outputs during test-time computation. This self-criticism serves as "privileged information" in the teacher context. Training then minimizes per-state divergence between the teacher's and student's denoising diffusion distributions, measured over the student's own sampling trajectories.
That means, as the abstract describes it, that the student model generates its own images during training (on-policy), and the teacher context — the same model, but with access to the self-criticism — acts as a reference point. Divergence minimization pulls the student's denoising process toward what the teacher "would have done" with the same information, without any external actor being needed to produce a label or a reward.
The numbers the authors report
The authors report the following results on Qwen-image-2512:
| Benchmark | Before | After |
|---|---|---|
| GenEval | 0.747 | 0.808 |
| GenEval2 Soft-TIFA | 32.97 | 35.53 |
These figures are author-reported and not independently verified. The publicly available basis for the paper is currently the abstract, and details about the training setup, reference models, and ablations cannot be confirmed beyond what the abstract provides. The paper's DOI is currently listed as "pending registration" on the arXiv page.
What the results suggest about the ceiling for self-improvement
The authors have also tried the method with stronger external critics, such as GPT5.6-Luna (this too is author-reported). Their finding is that multimodal models with strong judgment can expect a higher ceiling for self-improvement. In other words: the better the model is at evaluating the quality of its own work, the more it appears able to learn from itself. This is an interesting observation, but one that also depends on the judgment actually being good — otherwise the model risks reinforcing its own misinterpretations.
Mixed results on text rendering
The authors themselves are clear that the method's gains are not evenly distributed: "mixed text rendering results show that our self-improvements may not be uniform across different tasks," they write. Text rendering in images is a well-known weakness of diffusion-based models, and this suggests that self-criticism-based learning does not automatically transfer to all subtasks within image generation.
Open questions
Several key details remain to be verified from the full text:
- How large is the computational cost of on-policy self-distillation compared with conventional distillation?
- Which reference models were used in the GenEval comparison?
- Which ablations were run to isolate the effect of self-criticism as privileged information, versus other components of the training recipe?
- How does the method generalize to multimodal models other than Qwen-image-2512?
These questions are relevant to assessing whether UniEvo-VL represents a robust, replicable approach or a specific gain tied to this particular base model and these benchmarks.

