MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings
Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
Lens:Rethinking Training Efficiency for Foundational Text-to-Image Models
Scaling Properties of Text Conditioning in Visual Generation