A NEW PERSPECTIVE ON VISUAL REWARDS

Think before you score.

Thinking Reward Model for Visual Generation

Xuehai Bai1,4,*Zhenchen Tang2,*Yang Shi3,*,♠Dianyi Wang5

Tengfei Liu3Wanshun Su6Xuanyu Zhu3Ruohui Wang4Haiwen Diao7

Haotian Wang8,†Xiaoling Gu1,†Yuanxing Zhang3

1HDU2CASIA3PKU4SenseTime5FDU6NWPU7NTU8THU

* Equal Contribution♠ Project Lead† Corresponding Author

Complete paper teaser: Thinking Reward Model evaluates image editing and generation through adaptive rubrics, holistic analysis, and pointwise rewards; the right panel shows reward-model and downstream reinforcement-learning results.

A better score starts with a better question.

Visual reward models often map an image directly to a score, leaving the evaluation criteria implicit. TRM makes those criteria explicit: it generates rubrics for each individual case, inspects the candidate against them, and aggregates the evidence into a fine-grained reward. The same process supports both image generation and editing, while adapting what it checks to the task and the image.

Reason first. Reward with purpose.

A shared evaluation process across tasks. Criteria that adapt to every candidate.

01R

Determine

Formulate atomic rubrics around the task condition and candidate image.

WHAT SHOULD WE CHECK?
02J

Inspect

Evaluate the candidate against each rubric with an explicit Yes / No judgment.

WHAT DOES THE IMAGE SHOW?
03H

Aggregate

Summarize the evidence into assessments for each evaluation dimension.

HOW DO THE DETAILS ADD UP?
04s

Score

Produce a fine-grained pointwise reward grounded in the assessment.

HOW WELL DOES IT PERFORM?
PD-GRPO

Learn preferences. Keep the nuance.

Conventional pairwise optimization can push scores toward opposite extremes. Pairwise Dual-Group Relative Policy Optimization rewards sufficient separation, then stops rewarding further gap expansion.

The complete data construction and training pipeline is shown below, from structured annotation and supervised fine-tuning to preference optimization.

Complete training figure with unified data construction, two-stage annotation, supervised fine-tuning, and PD-GRPO preference optimization.
Data construction and training pipeline for the Thinking Reward Model.

Evaluation Results

Reward modeling benchmarks

Selected image generation reward-model benchmarks from the manuscript. Accuracy in percent; higher is better.
ModelParametersGenAI-T2I ↑MMRB2-T2I ↑
GPT-4.1 Proprietary—60.565.8
Gemini 3 Pro Proprietary—73.174.4
Qwen3-VL32B66.964.1
HPSv37B70.860.2
UnifiedReward7B67.959.8
RationalRewards8B69.864.2
Qwen3.5 · same-protocol baseline9B58.959.4
TRM (SFT)9B70.165.8
TRM (RL) OURS9B71.267.9

Selected baselines from the manuscript. Accuracy (%), higher is better. Pointwise models are evaluated over non-tied predictions.

Downstream task gains

Baseline TRM-guided RLΔ absolute score change

Downstream image generation results from the manuscript. Each cell shows baseline, TRM-guided RL, and absolute change. Higher scores are better.
ModelGenEval ↑DPG-Bench ↑TIIF-Short ↑TIIF-Long ↑
BAGEL0.860.89Change: +0.0385.0786.60Change: +1.5374.9180.68Change: +5.7775.6281.43Change: +5.81
FLUX.1-dev0.660.73Change: +0.0783.8485.28Change: +1.4470.8477.60Change: +6.7674.8281.38Change: +6.56
SD3.5-M0.660.72Change: +0.0684.2486.23Change: +1.9972.8678.04Change: +5.1872.8176.15Change: +3.34

The difference is in the details.

More faithful instructions. More precise compositions. Explore visual models before and after TRM-guided RL.

Image Generation

BAGEL · BASELINE & TRM-GUIDED RL
Complete BAGEL qualitative comparison figure, retaining all six baseline and TRM-guided pairs with their original prompts and labels.
Qualitative comparison of BAGEL before and after TRM-guided reinforcement learning.

Image Editing

SENSENOVA-U1.5 · SOURCE, BASELINE & AFTER RL
Complete SenseNova-U1.5 qualitative comparison figure, retaining all six examples, source images, baseline edits, optimized edits, and original Chinese and English instructions.
Qualitative comparison of SenseNova-U1.5 before and after RL fine-tuning with TRM.