we started out the summer trying to build the next generation of our VLM family, world class at classical image and video understanding. as we branched into tool calling capability, datasets which skew heavily towards traditional LLM domains like web search, we discovered some incredible cross task generalization: the model can control physical embodiments like drones and perform real world physical search tasks, demonstrating spatial awareness and navigational skills. this model is incredibly powerful, its fast, and you can throw basically any video, image, or real world agentic task at it. give it a try :-)
Today we're releasing Mk1.5: a new intelligence layer for embodied agents.
It flies drones, controls quadrupeds, powers smart glasses, tracks objects, searches the web, reasons visually, and dispatches its own sub-agents.
One model, no platform-specific retraining. 🧵