"Weโre using language as a crutch to help the deficiencies of our vision systems to learn good representations from images and video." Yann Lecun, Lex Fridman Podcast #416.
Sorry, this thread is very cool, but I still believe in the long run, building models with intrinsic understandings of the world requires a more vision-centric approach than modern LLMs.
You can just RL a coding model to paint with javascript btw