NEO-unify: Building Native Multimodal Unified Models End to End
Existing Multimodal AI Dilemma
For years, multimodal AI typically adopts a vision encoder (VE) to perceive and a variational autoencoder (VAE) to generate. Recent efforts seek to unify both with a