We benchmarked GPT-5.6 Sol (high), Grok 4.5 (high), and Claude Fable 5 on 24 browser tasks, run on Hyperbrowser with an identical harness,
3 trials per task.
Fable 5 led overall at 78%, Sol 75%, Grok 4.5 69% and each specialized: Fable was perfect on reading, Sol was best at navigation (94%), and Fable led on forms and logins (89%).
Full task set and harness are open source. The breakdown below ↓