When training a visual encoder with self-supervised learning, we know for a fact that using a decoder with a reconstruction loss doesn't work nearly as well as using a joint embedding architecture with feature prediction loss and a collapse prevention mechanism.
This paper from
@sainingxie's shop at NYU shows that *even* if you are interested in generating pixels (e.g. to produce pretty pictures with a diffusion transformer), it pays to include a feature prediction loss so that the internal representation of the decoder can predict features from a pre-trained visual encoder such as DINOv2.