Like Inkling, it's natively multimodal. It’s encoder-free, with audio and images processed jointly with text. It nearly matches Inkling across multimodal evals, and it can use Python to crop, zoom, and inspect images while reasoning over documents and charts.