Paper: UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
arXiv: 2512.21675
核心方向: MLLM / Perceptual Understanding / Image Reward
UniPercept 将 MLLM 从“语义理解”扩展到“感知层理解”,统一处理图像的美学、质量、结构与纹理,并通过 DAPT + Task-Aligned RL 训练,使模型同时具备评分(VR)和问答(VQA)能力。
最终还可以直接作为 T2I Reward Model。
传统 MLLM 擅长:
物体识别
场景理解
Caption
VQA
Visual Grounding
但不擅长:
这张图有多美?
图像质量有多高?
哪里存在 distortion?
结构是否规整?
纹理是否丰富?
即:
Semantic-level Understanding
↓
“图里有什么?”
Perceptual-level Understanding
↓
“图像看起来怎么样?”
论文认为目前 MLLM 在这一层存在明显不足。
论文首先构建统一的 Perceptual-Level Understanding Benchmark。
三个领域:
UniPercept-Bench
│
┌────────────┼────────────┐
↓ ↓ ↓
IAA IQA ISTA
Aesthetics Quality Structure & Texture