MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings

Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation

Lens:Rethinking Training Efficiency for Foundational Text-to-Image Models

Scaling Properties of Text Conditioning in Visual Generation