LongCat-Next: Lexicalizing Modalities as Discrete Tokens(美团)

Beyond Language Modeling:An Exploration of Multimodal Pretraining

UNDERSTANDING VS. GENERATION: NAVIGATING OPTIMIZATION DILEMMA IN MULTIMODAL MODELS

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings

TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models

UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer

**MiniT2I:A Minimalist Baseline for Text-to-Image Generation

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer

ERNIE 5.0 Technical Report(百度)

NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation