Reasoning ability and quality of judgement are different axes for humans as well.
There's people who are very good at constrained reasoning, yet find poor answers for open-ended questions.
And some people can make reasonable choices while not being able to pass a logic test.
With LLMs, we like to think of "smart" and "dumb" as one axis (because we think of humans this way).
I'd like to argue against this framing. Instead, try to think of "smart" and "dumb" as two different axes. A model can be incredibly smart AND dumb at the same time.
For a good example of this, look at a Gemini model. They are incredibly smart: you can bench them and see the capability. The amount of knowledge Google bakes into their models is incredible. Yet when you ask it to do work, the amount of stupid things the models will do is similarly incredible. They are "smart" and "dumb" at the same time.
Similarly, look at a model like Fable 5.1. It is not quite as smart as Astra, but it is significantly less dumb. Astra is still the smartest model available today, but it is also one of the dumbest, regularly doing things that make literally no sense whatsoever.
Finding the right balance of both "smart" and "dumb" is a challenge that everyone needs to figure out, from users to researchers. If we stop thinking of it as either-or and instead realize both can exist together, understanding model behavior gets much easier.